Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,342 characters · 17 sections · 61 citation commands
\setcounter{page}{1} \pagenumbering{arabic}
Designing optimal education and labor markets, policy-makers face an increasingly important yet difficult challenge of assessing critical aspects of occupations. For instance, Autor2003,Autor2006,Acemoglu2011,Autor2013skill,Autor2013 find the occupational degree of abstract, routine, and manual job tasks to be important in explaining changes in the labor market; Goldin2014 examines the role of various occupational characteristics in explaining the gender wage gap; Deming2017 consider the degree of social skills in occupations; Frey2017 examine the susceptibility of occupations to computerization; Brynjolfsson2018a estimate an occupational suitability index for machine learning (ML); Bowen2018 examine the “greenness” of occupations; and Dingel2020 and Mongey2021 construct measures of the occupational feasibility of working at home and exposure to social distancing, respectively, to study the economic impact of social distancing in the perspective of COVID-19.
Common to these seminal studies is the need to measure fundamental characteristics of occupations. But all existing studies use case-by-case approaches that are deeply difficult to generalize and validate. Most on them rely on detailed data from the Occupational Information Network (O*NET) that provides measures of hundreds of occupational features. Although the richness of the data certainly allows for novel contributions, an inherent risk of overfitting lurks in the dark. That is, researchers might be tempted to select exactly those features of the data that makes the analysis fit the narrative. The real difficulty is, however, that it may be borderline impossible to validate and verify truly novel measures of occupational characteristics. Another challenge is that every new measure requires a new approach, method, study, etc., which is surely laborious.
We propose occ2vec, a principal approach to representing occupations as high-dimensional vectors, which can then be used in matching, comparative studies, predictive and causal modeling, and other economic areas of interest. Specifically, we demonstrate how the high-dimensional occupation vectors can be used to score occupations on any definable target characteristic, for instance, the occupational degree of “greenness”. At its core, our approach essentially transforms every occupation into a high-dimensional vector for which reason we call it occ2vec. The only input needed is an objective and reliable textual definition of the target characteristic. For instance, the U.S. Bureau of Labor Statistics defines green jobs as Definition (ref):
Using O*NET, we rely on 244 occupational attributes, 873 occupation descriptions, and 16,804 occupation-specific tasks to learn a vector representation of each occupation, leveraging natural language processing (NLP).\footnote{In principal, any database containing detailed information on the universe of occupations could be used. To our knowledge, O*NET provides the most comprehensive information on occupations. In addition, O*NET provides both textual descriptions and numerical scores, which is necessary for our validation strategy. This will be thoroughly explained in Section (ref). An alternative source of occupation data would be the European Skills, Competences, Qualifications and Occupations (ESCO) data, which can be found \href{https://esco.ec.europa.eu/en}{here}. We are not familiar with any papers using the ESCO data, and thus we leave this for further research.} Once all occupations are embedded as high-dimensional vectors, we consider a given target characteristic of interest and assign to it a vector in the same vector space as the occupations. This allows us to estimate the occupational degree of the target characteristic as the (standardized) cosine similarity between its vector and any occupation vector.
The advantage of this framework is that it is fully data-driven and universal, only requiring the user to provide a reliable and objective definition of the target characteristic. Hence, this is easily extendable to genuinely novel attributes of occupations. In principle, one could extend the framework to default to using an encyclopedia, e.g., Wikipedia, and then the approach would be truly automated. Another advantage is how easy it is to validate the framework in contrast to existing approaches that do not consider any ground truth. We validate occ2vec in two ways. First, we visually inspect the occupation vectors by compressing them into two-dimensional vectors using principal component analysis (PCA) and $t$-distributed stochastic neighbor embedding ($t$-SNE) vandermaaten08a. We plot all occupations on the two dimensions and show that occupations cluster according to major occupational groups and educational requirements. This indicates that even after compressing the information from the high-dimensional vectors into two dimensions, the low-dimensional occupation vectors still precisely capture differences between occupations. Second, we estimate the occupational degree of all of the 244 occupational attributes from O*NET and compare our estimates to the original O*NET scores. We find strong evidence that our estimates coincide with the original O*NET scores both between and within occupations and take this altogether as evidence that occ2vec produces high-quality occupation vectors that capture essential features of the occupations. These vectors can then be used to match occupations more precisely, act as control covariates in regressions, etc. Specifically, we demonstrate how the vectors can be used to score occupations on novel characteristics.
Once the quality of the occupation vectors has been scrutinized and confirmed, we showcase the framework using several applications, which we divide into estimating well-studied versus novel characteristics of occupations.
First, we revisit the popular task measures (abstract, manual, and routine tasks, respectively) by Autor2003,Autor2006,Acemoglu2011,Autor2013skill,Autor2013 and the innovative measures of artificial intelligence (AI) by Felten2018, Brynjolfsson2018a, and Webb2019. This illustrates how one could have used our framework had the other case-specific approaches not been invented. Note carefully that this is not another validation exercise, except under the assumption that the existing measures represent the ground truth. We feel more confident making this assumption for the original O*NET scores.
We document that our estimates of each task measure coincide well with the original measures, and we find that performing a high amount of routine tasks is on average associated with lower educational attainment and lower wages. We find the same education and wage patterns for occupations that are characterized by performing manual tasks, whereas the opposite holds for abstract tasks.
We find that our estimates of AI exposure generally behave as a mix of the measures of Felten2018 and Webb2019 with no clear relationship to that of Brynjolfsson2018a. The close similarity of Felten2018 and Webb2019 in contrast to Brynjolfsson2018a was first noted by Acemoglu2022ai. Occupations that have a high degree of AI have similar profiles as those that score high on abstract tasks. That is, these occupations generally have higher educational requirements and enjoy higher wages.
Second, we investigate the occupational degree of two novel characteristics, namely charisma and emotional intelligence (EQ). We use definitions from psychological outlets, e.g., the American Psychological Association. Both characteristics are intrinsically interesting and have been found to play a fundamental role in leadership skills and social affability. At the outset, charisma and EQ seem conceptually closely related, and we do find similarities empirically. The occupations that score high on both charisma and EQ are often found within community and social service, educational instruction, and arts, design, entertainment, sports, and media occupations. as well as educational instruction. For charisma, another frequent occupational group is sales. The occupations with high degrees of charisma and EQ also share similar wage profiles. Particularly, our estimates suggest occupations in the upper part of the wage distribution score high on both charisma and EQ. One difference appears in the lower part of the wage distribution, where occupations tend to score high on charisma but not on EQ. Another difference that we detect at the aggregate level enters between occupations that require master's degree versus doctoral degrees. Essentially, these occupations score similarly on charisma, whereas there is a dip in the level of EQ for occupations that require doctorates.
The rest of the paper is organized as follows. Section (ref) presents the data used in occ2vec, whereas Section (ref) introduces the framework and our NLP method of choice. Section (ref) validates our framework and Section (ref) considers our four applications on task measures, AI, charisma, and EQ, respectively. Section (ref) concludes. Appendices to data, method, and applications are found in Appendix (ref), (ref), and (ref)--(ref), respectively. Throughout the paper we will be using the notation $\left|A\right|$ for the cardinality of a generic set $A$ and $\left[a,b\right]=\left\{ z\in\mathbb{Z}\vert a\leq z\leq b\right\}$ denotes the set of integers between $a$ and $b$ with $\left[a\right]$ simply meaning $\left\{ 1,\ldots,a\right\}$.
In this section, we describe the data used to construct the occupation vectors. Starting with the data has two advantages. First, in our experience, it is easier to understand the framework with a specific source of data in mind, especially for readers less versed in NLP. Second, although any detailed occupational database could be used, O*NET provides the most comprehensive information to our knowledge. Specifically, the validation strategy requires the source of occupational data to contain both textual descriptions and numerical scores for descriptors of the entire universe of occupations, which we have not come across elsewhere.\footnote{Specifically, the O*NET\textsuperscript \textregistered Content Model, which can be found \href{https://www.onetcenter.org/content.html}{here}. We use O*NET version 26.3, which can be found \href{https://www.onetcenter.org/dictionary/26.3/excel/}{here}.}
O*NET keeps track of hundreds of descriptors for each occupation and we use the 873 unique occupations available. Each occupation is described textually by a general description (e.g., writers generally originate and prepare written material, such as scripts, stories, advertisements, and other material) and by detailed tasks (e.g., bakers place dough in pans, molds, or on sheets), as well as measured numerically on several attributes (e.g., how important is critical thinking for surgeons). We denote the three sources of occupational information (i.e., descriptions, tasks, and attributes) collectively as occupational descriptors. We highlight this distinction because the majority of the existing approaches is only suited for using the numerical scores of the attributes, whereas our method incorporates purely textual data as the tasks and descriptions as well.\footnote{One important exception is Webb2019, who uses the tasks but not the attributes. To the best of our knowledge, Webb2019 cannot be generalized to using the attributes as well.} Note that attributes are also defined textually and that tasks are also associated with a weight (i.e., a numerical score). Thus, all the occupational descriptors are both expressed as text and associated with an occupation-specific numerical score. For instance, the definitions of the attributes oral comprehension and depth perception follow from Table (ref). These textual definitions naturally apply to all occupations, but each occupation is also assigned a specific score between zero and one for all attributes. For instance, being an astronomer is associated with a score of 0.74 on oral comprehension and of 0.25 on depth perception.
Similar, Table (ref) exemplifies an occupation description as well as a few tasks for astronomers. Additionally, the descriptors are associated with a weight.
The occupational descriptors represent ten categories; namely, description, tasks, abilities, interests, work values, work styles, \textit{skills}, \textit{knowledge}, \textit{work activities}, and \textit{work context}.\footnote{Technically, the ten categories further represent four broader types of information about work; worker characteristics, worker requirements, occupational requirements, and occupation-specific information. Specifically, worker characteristics include abilities, interests, work values, and work styles; worker requirements include skills and knowledge; occupational requirements include work activities and work context; and occupation-specific information include description and tasks. In this paper, we focus on the ten categories and we do not distinguish further between the four broader types of occupational information.} The total number of occupational descriptors amounts to 17,048, which includes 873 occupation descriptions (one for each occupation), 16,804 unique tasks (some tasks are occupation-specific, while others are performed by a few occupations, leading to on average 20 tasks per occupation), and 244 attributes (common across occupations but with occupation-specific weights). Figure (ref) shows the structure of the data.
We are exhaustive in the selection of occupational descriptors in the sense that we include all those for which meaningful textual descriptions are included.\footnote{For instance, the ability stamina is defined by the ability to exert yourself physically over long periods of time without getting winded or out of breath, whereas the technological skill electronic mail software (e.g., Microsoft Outlook) is not further defined.} The occupation-specific weight that is associated with each descriptor is constructed based on at least one scale, e.g., importance, level, etc., and the realized value of the scale differs by occupations. The definition of each scale along with the set of categories that each scale applies to are presented in Table (ref) in Appendix (ref). Each scale has a minimum and maximum value and to enable comparison, all scales have been standardized to ranging from 0 to 1 following O*NET guidelines. In the case of multiple scales for a specific descriptor, we take a uniform average of the standardized scales.
In this section, we present our principal approach to measuring any definable target characteristic of an occupation by quantifying the similarity between two vectors that represent the occupation and the target characteristic, respectively. The occupation vectors are text embeddings (more intuition below) generated by our preferred NLP method that heavily rely on the plethora of occupation-specific information provided by the O*NET, and thus, our approach essentially transforms every occupation into a high-dimensional vector that captures the tasks and attributes of the occupation. Similar, the vector of the target characteristic is a text embedding of its definition that can in principle be provided by any objective and reliable source of information, e.g., international standards or dictionaries. The characteristic of interest, say the “greenness” of occupations, is thereby also transformed into a high-dimensional vector, which allow us to compute a measure of similarity of the two vectors; one vector representing a particular occupation and one vector representing the characteristic.
Before deep-diving into details, we provide some intuition of the text embeddings that are the central building blocks of occ2vec. Overall speaking, a text embedding is a real-valued vector that represents a mathematical analogue of the piece of text in question such as words, sentences, paragraphs, etc. The vector encodes the meaning of the text such that the texts that are closer in the vector space are expected to be similar in meaning. In a highly-simplified one-dimensional world, this would mean that the occupation “actuary” could have a mathematical representation of, say, the number 2 and “mathematician” the number 3. This means that linguistically “actuary” and “mathematician” are similar in the sense that their tasks and attributes overlap to a high degree. In contrast, imagine “actors” and “dancers” would be represented by, say, 11 and 14, respectively, meaning that actors and dancers are more similar to each other than to either actuaries or mathematicians.\footnote{In modern applications of text embeddings, the dimensions are several hundreds and we use $p=1024$.} Thus, the process of mapping pieces of texts to real numbers leads to the creation of embeddings that capture the meaning of the text. The benefit of using text embeddings is that it allows texts with similar meanings to have similar vector representations (as measured by high cosine similarity).
Embedding more than 17,000 textual descriptors, plenty of NLP methods exist. Common to all is the assumption of the Distributional Hypothesis due to Harris1954, claiming that words that occur in the same contexts tend to have similar meanings and also popularized by Firth1957 as “a word is characterized by the company it keeps”. Some of the natural choices to consider are Word2Vec Mikolov2013EfficientEO, its document version Doc2Vec/Paragraph2Vec Quoc2014, GloVe Pennington2014, Fasttext Bojanowski2016,Joulin2016a,Joulin2016b, or BERT Devlin2018. In this paper, we use an optimized version of BERT called Sentence-RoBERTa reimers2019,Liu2019, and thus our foundation is BERT. Our framework is not limited to specific embedding techniques, and our findings are robust to the choice of NLP methods. We will not review all details of BERT/RoBERTa as this is out of scope. Instead, we provide an intuitive explanation of BERT/RoBERTa and refer readers to Appendix (ref) for more details or Rogers2021 for an excellent review.
Let $T$ be a given piece of text (e.g., a sentence or a paragraph) from corpus $\mathcal{T}$. Each $T$ is essentially a sequence of tokens (e.g., words or subwords). Further, let $D\in\mathbb{R}^{p}$ be an embedding of $T$, i.e., a $p$-dimensional vector that mathematically represents the meaning of $T$. At least conceptually, the assumption underlying all approaches to text embeddings is that there exists a map $f:T\rightarrow D$ approximating the relationship between $T$ and $D$. Any text embedding model may be viewed as a nonparametric estimate $\hat{f}$ of the map $f$. The challenge is that the true embeddings are latent, and thus not observed. The various methods consider different approaches to circumventing this, but the most common routine to learn $f$ is to mask a random sample of tokens and then train the model to produce embeddings by minimizing the cross-entropy loss from predicting the masked tokens given the estimated embeddings, which is essentially a cloze procedure Taylor_1953. The training process encodes a lot of semantics and syntactic information about the language by training the model on massive amounts of unlabeled textual data drawn from the web. The objective is to detect generic linguistic patterns in the text.
Index occupations by $i$ for $i\in\left[n\right]$ and assume access to data $\left\{ \left(W_{j},T_{j}\right):j\in\left[d\right]\right\}$, where $T_{j} \in\mathcal{T}$ is the textual definition of the $j$th occupational descriptor, $\mathcal{T}$ represents the universe of unique descriptor definitions, and $W_{j}\in\left[0,1\right]^{n}$ is an $n$-dimensional vector of occupation weights on the $j$th descriptor such that $w_{i,j}$ represents the importance of descriptor $j$ for occupation $i$. Let $\mathbf{W}=\left(W_{1}\cdots W_{d}\right)\in\left[0,1\right]^{n\times d}$ denote the matrix of concatenated weights.\footnote{We normalize $\mathbf{W}$ such that each $\mathbf{W}_{i,\cdot}$ sum to unity for $i\in\left[n\right]$. We further normalize $\mathbf{W}$ such that each of the ten O*NET categories contribute equally, that is, weights representing the same category sum to one over the number of categories.} In principle, all occupations could be represented by any descriptor but in practice many weights are zero because the descriptors contain tasks only performed by a subset of occupations.
The goal of the NLP algorithm of choice is to turn each descriptor definition $T_{j}$ into a descriptor embedding $D_{j}\in\mathbb{R}^{p}$, where $p$ represents the dimensionality of the embeddings, leading to an embedding matrix $\mathbf{D}=\left(D_{1}\cdots D_{d}\right)\in\mathbb{R}^{d\times p}$. The occupation embeddings are then constructed as a weighted average of the descriptor embeddings via (ref), i.e.
where $\mathbf{X}=\left(X_{1} \cdots X{}_{n}\right)\in\mathbb{R}^{n\times p}$, and $X_i \in\mathbb{R}^{p}$ represents occupation $i$ as a $p$-dimensional vector. Note that $p$ is typically several hundreds and we use $p=1024$. Using this representation of occupations, we think of a given occupation as a weighted average of its descriptors (attributes, tasks, and title), where the weights are governed by the combined importance, relevance, and frequency of the descriptor.
Tailoring occ2vec to occupational data, we develop a novel fine-tuning step.\footnote{Since we only have access to 873 occupations via O*NET, we cannot fine-tune the embeddings in the standard way of formulating a deep neural network to classify the occupations.} Imagine access to a new, unseen target occupational characteristic, $T_0$, and an $n$-dimensional vector $Y_0$ of occupational scores on $T_0$. Similar to the construction of $\mathbf{D}$, we use our NLP algorithm to embed $T_0$ into a $p$-dimensional vector $D_0$. Our goal is then to fine-tune the occupation embeddings, $\mathbf{X}$, such that the cosine similarity between the occupation embeddings and the target characteristic embedding predicts the score, $Y_0$, as closely as possible. The cosine similarity between a particular occupation embedding $X_i$ and the target characteristic embedding $D_0$ is given by
which simplifies to the dot product, $S_{\cos}\left(X_{i},D_{0}\right)=X_{i}\cdot D_{0}$, because both embeddings are normalized to unit length. Our approach to fine-tuning the occupation embeddings is to estimate a $\left(p\times p\right)$-dimensional rotation matrix $\mathbf{R}$ that solves
Estimating $\mathbf{R}$ via Eq. (ref), it is most appropriate\footnote{Technically, we avoid severely sparse matrices by focusing on common descriptors.} to consider descriptors that are shared across all occupations, namely the attributes in the case of O*NET. For this reason, let $\mathcal{D}$ denote the entire set of descriptors and let $\mathcal{A}$ denote the (smaller) set of attributes that are shared among occupations, such that $\mathcal{D}\setminus\mathcal{A}$ represents tasks and job titles. Then, let $\mathbf{Y}=\mathbf{W}_{\cdot,\left\{ j:j\in\mathcal{A}\right\} }\in\left[0,1\right]^{n\times a}$ be the subset of columns in $\mathbf{W}$ that represents the attributes among all descriptors, where $a=\left|\mathcal{A}\right|$.\footnote{Recall that descriptors include attributes (shared among occupations), tasks (unique to subset of occupations), and titles (unique per occupation).} We interpret each entry in $\mathbf{Y}$, namely $y_{i,j}$ for $i\in\left[n\right]$ and $j\in \mathcal{A}$, as representing the score of occupation $i$ on attribute $j$. Similarly, let $\mathbf{A}=\mathbf{D}_{\left\{ i:i\in\mathcal{A}\right\} ,\cdot}\in\mathbb{R}^{a\times p}$ denote the subset of rows in $\mathbf{D}$ that represent the attributes, such that $\mathbf{A}_{j,\cdot}$ represents the $p$-dimensional vector embedding of attribute $j$. The empirical matrix analog to (ref) then reads
where subscript UR abbreviates unrestricted, and $\left\Vert \cdot \right\Vert _{F}$ denotes the Frobenius norm. The optimization problem in Eq. (ref) without any constraints has an analytic solution given by
provided that $\mathbf{X}$ and $\mathbf{A}$ have full column rank such that the inverses exist.
Empirically, we find, however, that using Eq. (ref) rather than imposing additional constraints leads to sub-optimal performance on unseen data.\footnote{Essentially, setting $\mathbf{R}=\mathbf{I}_{p}$ performs comparably to $\mathbf{R}_\text{UR}$ on data not used in the estimation of $\mathbf{R}_\text{UR}$} This is most likely due to overfitting as the unconstrained optimization problem is highly overparametrized. Thus, we experiment with three additional constraints, namely $\mathbf{R}_{k,k'\in\left[p\right]:k,k'\in k\neq k'}=0$ (ensuring that $\mathbf{R}$ is diagonal), $\mathbf{R}_{k,k'\in\left[p\right]}\geq0$ (ensuring that all entries in $\mathbf{R}$ are non-negative), and $\sum_{k}\sum_{k'}\mathbf{R}_{k,k'}\leq c$, where $c$ is some constant (ensuring that the sum of entries in $\mathbf{R}$ is constrained to avoid too much flexibility). We always enforce $\mathbf{R}$ to be diagonal and refer to the rotation matrices stemming from each constraint as $\mathbf{R}_\text{DIAG}$, $\mathbf{R}_\text{DIAG+NONNEG}$, and $\mathbf{R}_\text{DIAG+SUM}$, respectively. To no-tuning benchmark is $\mathbf{R}=\mathbf{I}_{p}$, where $\mathbf{I}_{p}$ is the $p$-dimensional identity matrix, which we denote $\mathbf{R}_\text{IDENTITY}$.
To avoid overfitting, we never re-use the same data to estimate the occupation vectors in $\mathbf{X}$ and the rotation matrix $\mathbf{R}$, or to evaluate the performance. Hence, we randomly split the set of attributes multiple times into three equal-sized sets without replacement such that $\mathcal{A}=\mathcal{A}_{1}\cup\mathcal{A}_{2}\cup\mathcal{A}_{3}$ and $\mathcal{A}_{h}\cap\mathcal{A}_{g}=\emptyset\;\forall h\neq g$. Then, define $\mathbf{Y}^{g}=\mathbf{Y}_{\cdot,\left\{ j:j\in\mathcal{A}_{g}\right\} }$ for $g=1,2,3$, such that the concatenation $\tilde{\mathbf{Y}}=\left(\mathbf{Y}^{1},\mathbf{Y}^{2},\mathbf{Y}^{3}\right)$ is simply a reshuffled version of $\mathbf{Y}$. Similarly, define $\mathbf{A}^{g}=\mathbf{A}_{\cdot,\left\{ j:j\in\mathcal{A}_{g}\right\} }$ for $g=1,2,3$. Last, define $\mathbf{W}^{g}=\mathbf{W}_{\cdot,\left\{ j:j\notin\cup_{h\neq g}\mathcal{A}_{h}\right\}}$ and $\mathbf{D}^{g}=\mathbf{D}_{\left\{ i:i\notin\cup_{h\neq g}\mathcal{A}_{h}\right\} ,\cdot}$ for $g=1,2,3$. Having the definitions in place, we compare the various strategies for obtaining an estimate of $\mathbf{R}$ by the following procedure:
We re-do the above procedure five times (with a random split of attributes each time) and report the performance relative to the benchmark in Table (ref).
Since our objective is to generalize the performance of our occupation embeddings to unseen target characteristics, the most interesting rows are in the bottom of Table (ref), representing the out-of-sample performance. The best performing strategy on both loss functions is $\mathbf{R}_\text{DIAG+NONNEG}$, where the rotation matrix is forced to be diagonal and have non-negative entries. The gain relative to the benchmark is almost 16% in terms of the Frobenius norm, which we optimize with respect to. Note also from Table (ref) how especially $\mathbf{R}_\text{UR}$ is prone to overfitting since the in-sample performance is essentially perfect but the out-of-sample performance is actually marginally worse than the benchmark. Thus, we overwrite the occupation vectors and shall henceforth use $\mathbf{X}\coloneqq\mathbf{X}\mathbf{R}_{\text{DIAG+NONNEG}}$.
In this section, we validate occ2vec by performing two exercises. The first exercise is exploratory and visual, where we show that occupation vectors can be well separated by major occupational groups and required educational attainment. The second exercise is more formal and statistical, where we use various metrics to assess the performance of our estimates of the same occupational attributes that O*NET provides. Note that if one is interested in characteristics that appear in the O*NET database (any of the 244 occupational attributes), there is no reason to use our approach. In this case, one should use the O*NET scores directly. Our method is relevant when one is considering novel characteristics that are not included in O*NET. Taking the O*NET scores as the ground truth simply allows for an transparent validation of our approach.
Having outlined how we quantify an occupation by transforming it to a high-dimensional vector, we illustrate the occupations in a two-dimensional space and color them by major occupational groups and educational requirements in Figure (ref) and (ref), respectively. Reducing the 1024 dimensions of the embeddings to only two dimensions, we first use PCA to reduce the number of dimensions to 50 and then we use $t$-SNE vandermaaten08a to further reduce the dimensionality to two.\footnote{This is a popular tool to visualize high-dimensional data, which works by converting similarities between data points to joint probabilities and minimizing the Kullback-Leibler divergence between the joint probabilities of the low-dimensional embedding and the high-dimensional data. }
From Figure (ref), it follows that the 873 occupations can robustly be separated into the 22 major occupational groups as the occupations tend to cluster into these groups (of the same color). Note that the different colors have no interpretation as groups are sorted alphabetically. For instance, both production occupations in purple, construction and extraction occupations in green, and installation, maintenance, and repair occupations in blue can be found in the LHS, and despite the differences in colors, these groups represent a cluster related to physical and manual occupations.
Likewise, Figure (ref) indicates that educational requirement is an important separator of occupations. Especially the two extremes, i.e., high school diploma and doctoral degree, occupy two different positions in the two-dimensional space. Note that the colors do have an interpretation is this figure, because we can order educations by level.
The two figures support the idea that occupations are quantifiable and separable and this is despite the fact that we compress all the information in the high-dimensional space to only two dimensions. When comparing the figures, note that they would be identical without the coloring because we only perform the dimension reduction once.
Recall that O*NET provides 244 occupational attributes for each of the 873 occupations. Principally, we may estimate the occupational degree of each of these attributes the same way as we would for any target characteristic of interest. That is, we compare the occupation embeddings to each of the attribute embeddings, thereby quantifying the degree of each attribute as the cosine similarity. This gives us 244 estimates of occupational attributes for each of the 873 occupations, totaling 213,012 estimates we can readily use to validate our framework assuming O*NET to be the the ground truth. This is possible because O*NET provides both numerical scores as well as textual definitions for all attributes.\footnote{Note that we cannot do this for the 873 occupation descriptions nor the 16,804 occupation-specific tasks both because these descriptors not comparable across occupations.}
\paragraph*{Between and within occupations} First, we assess the ability to correctly estimate an occupational attribute by considering correlations between and within occupations. The reason we care about a high correlation between occupations is that for a specific attribute, we should be able to rank the occupations accordingly (e.g., how do judges compare to carpenters on critical thinking?). To assess between-occupations validity, we use the 213,012 attribute estimates and compute the Spearman correlation coefficient between the O*NET score and our estimate by attribute, yielding 244 correlation estimates.
A high correlation within occupations is equally important because for a specific occupation, we should be able to rank multiple attributes (e.g., how does static strength compare to dynamic strength for firefighters?). To assess within-occupations validity, we again use the 213,012 attribute estimates and compute the Spearman correlation coefficient but this time by occupation, yielding 873 correlation estimates.
Obtaining two sets of correlations, we determine if each mean correlation differs significantly from zero using classic $t$-tests. In additional, we test if each mean correlation differs significantly from hypothesized values ordered from 1% to 99% and report the first postulated correlation for which we fail to reject the null hypothesis of no significant difference at the 5% level. The results follow from Table (ref), where Table (ref) shows the between-occupations tests and Table (ref) shows the within-occupations tests.
Both types of validity tests strongly reject the null hypothesis of zero correlation. In fact, it is not until reaching a postulated correlation coefficient 53% between occupations and 77% within occupations that we fail to reject the null hypothesis.\footnote{These results indicate that our framework is relatively more capable of ranking attributes for a specific occupation.} Together, the validation exercises suggest a significant non-zero mean correlation between the O*NET scores and our estimates.
\paragraph*{Overall explainability} As a second validation exercise, we run various regressions of the O*NET scores on our estimates (and a constant), using the entire set of 213,012 attribute estimates. The specifications we consider are the eight combinations that may be generated using occupation, attribute, and category dummies, respectively. That is, the baseline specification includes no control dummies, whereas the full specification includes all three sets of dummies. In Table (ref), we report the results from all eight regressions. The first row shows the coefficient on our measure, the standard error in parentheses, and the $t$-statistic in brackets. Rows 2-4 details the specification and the last two rows show the adjusted $R^2$ and the number of observations, respectively.
From Table (ref), the coefficient under scrutiny is highly significant and strongly robust across all specifications, with $t$-statistics above 300 in all specifications. In addition, the adjusted $R^2$ in the baseline specification reaches 59%. This supports the consistency, reliability, and validity of our proposed framework.
In this section, we estimate the occupational degree of both some well-known characteristics (task measures and AI measures, respectively), and some novel characteristics (charisma and EQ, respectively). We show how the occ2vec framework ranks occupations according to the characteristics and how our estimates correlate with various occupational statistics, e.g., educational requirements and wage. All occupational statistics are sourced from the U.S. Bureau of Labor Statistics. For the well-known characteristics, we also assess the similarity of our estimates to those already established in the literature.
We revisit the seminal work of Autor2003,Autor2006,Acemoglu2011,Autor2013skill,Autor2013 to study tasks measures, namely abstract, manual, and routine tasks. In Table (ref) in Appendix (ref), we outline the individual O*NET attributes that comprise the tasks measures and refer to Acemoglu2011 for details on precisely how to construct them. We will mainly focus on the degree of routine tasks and refer to Appendix (ref) for similar analyses on the degree of abstract and manual tasks, respectively.
\paragraph*{Routine tasks} We start by defining routine tasks in Definition (ref), relying heavily on Acemoglu2011.\footnote{We thank Daron Acemoglu for providing feedback on this definition.} Recall how we feed this definition into our NLP method, which returns an embedding in the same vector space as the occupations. We then measure the occupational degree of routine tasks by the cosine similarity between each occupational embedding and the embedding of routine tasks.
In Table (ref), we list the top and bottom 10 occupations on the degree of routine tasks. The results are as expected and the NLP algorithm even picks up the specific examples in Definition (ref) such as clerical work (e.g., rank 2) and bookkeeping (e.g., rank 9). All top 10 occupations are truly characterized by accomplishing codifiable tasks that can be specified as a series of instructions to be executed by a machine. In stark contrast, the tasks of choreographers or psychiatrists are by no means codifiable.
Given the ability of the framework to separate occupations by major occupational groups and educational requirements, we show the degree of routine tasks by these two attributes in Figure (ref) and (ref), respectively. Figure (ref) confirms that occupations that accomplish a higher degree of routine tasks are to be found in office and administrative support occupations, where the opposite can be found in community and social service occupations as well as educational instruction and library occupations. Figure (ref) further shows that routine tasks are characteristic of occupations that require low-to-medium educational attainment, where occupations that require bachelor's, master's or doctoral degrees rarely involve routine tasks. In fact, the occupational degree of charisma and the educational requirement are inversely, monotonically related.
We next highlight how our estimates largely coincide with the original measures by Autor2003. Essentially, we estimate a smoothed polynomial regression of both measures of routine tasks for each occupation against its rank in the wage distribution. The result is shown in Figure (ref). Overall, the two approaches agree nearly perfectly on the relationship between the degree of routine tasks and the wage rank; occupations in the upper part of the wage distribution tend to be characterized by smaller amounts of routine tasks, whereas the opposite holds for the lower part of the wage distribution where mainly occupations with many routine tasks reside. A trend that has been evidenced by many (see, e.g., Acemoglu2011).
As a final comparison, we compute the Spearman correlation coefficients between our estimates of abstract, manual, and routine tasks and the original ones. The correlation coefficients are 0.40, 0.84, and 0.64, respectively. Thus, the degree of association between the two approaches appears strong and significant, particularly for manual and routine tasks.
Turning to the occupational degree of AI, we use the Definitions (ref)--(ref) presented in Appendix (ref). Similar, all figures and tables for this application is deferred to Appendix (ref). We start by considering the top and bottom 10 occupations on degree of AI from Table (ref). In top 10, we find e.g., business intelligence analysts and statisticians, which is aligned with our expectation as these occupations involve tasks that require various forms of e.g., predictive and prescriptive analytics, which are integral parts of AI. In the bottom 10, we find e.g., manicurists, pedicurists, and stonemasons.
Considering the degree of AI by major occupational groups and educational requirement, we show Figure (ref) and Figure (ref), respectively. As expected, the top major occupational groups computer and mathematical occupations, which coincides with the top occupations from Table (ref). The educational requirements tend to be higher for occupations with high degrees of AI, which makes sense as AI is difficult to skill to acquire, but the association is not monotonically increasing in education. In fact, occupations requiring bachelor's or master's degrees tend to have a higher exposure to AI than occupations that require doctoral degrees.
Next, we compare our estimates of occupational degree of AI to the already established measures, namely those of Felten2018, Brynjolfsson2018a, and Webb2019. In Figure (ref), we repeat the previous analysis and estimate a smoothed polynomial of the measures for each occupation against its rank in the wage distribution. Our measure of the occupational degree of AI is to be found in the middle between the measures of Felten2018 and Webb2019, meaning that its relationship to occupational wage and employment growth is a mix of those of these well-established measures. Interestingly, our measure of occupational AI tends to agree relatively more with the one of Felten2018 for the high-wage occupations and agree relatively more with the one of Webb2019 for the low-wage occupations. In fact, the Spearman correlation coefficient between our measure and those two is 0.35 and 0.2, respectively, as shown in Table (ref). Specifically, occupations receiving higher wages tend to have a higher degree of AI compared to the occupations that receive less. The differences between the AI--wage profiles shown in Figure (ref), however, suggest that the three established measures are picking up different aspects of AI. A fact that is already found by Acemoglu2022ai, who also find that the measures of Felten2018 and Webb2019 tend to coincide and be different than the one of Brynjolfsson2018a.
The second part of this section considers two novel characteristics of occupations, namely the degree of charisma and EQ. To economize on space, all figures and tables for the EQ application are deferred to Appendix (ref). Definitions are likewise found in Appendix (ref) and (ref), respectively.
The first systematic treatment of charisma is due to Weber1947 and popularized by Dow1969. Charisma is at the very center of effective leadership Avolio2013 and is continuously being studied (see, e.g., Hippel2016). It has been recognized as one of the main explanations for why certain leaders, i.e., charismatic leaders, develop emotional attachment with followers and other leaders that eventually foster performance that surpass expectations. Likewise, charisma is associated with being considerate, inspirational, visionary and intellectually stimulating (for a review, see, e.g., Spencer_1973 or Turner_2003). One definition of charisma follows from Definition (ref), which is part of the definitions we use to construct the charisma embedding.\footnote{The other definitions that are included in the construction of the embedding of charisma follow from Definitions (ref)--(ref) in Appendix (ref).}
We begin our analysis by highlighting the top and bottom 10 occupations on the degree of charisma in Table (ref). Public Relations specialists, actors, and fundraisers are among the top occupations that associate with charisma, because those occupations require both the ability to catch the attention of the audience and the knowledge of group behavior and dynamics, societal trends and influences. Similar for school psychologists who must be knowledgeable of human behavior and be capable of assessing individual differences in personality and interests. An interesting yet obvious finding is that directors within religious activities and education are among the top occupations as these directors must be able to speak to the public and gain followers by promoting the religious education or activities and teaching their religion's doctrines. Generally, occupations within community and social services as well as sales score high on charisma. On the other side of the spectrum, we mostly find occupations within production, construction, farming, fishing, and forestry. These occupations have less of a need to be able to inspire and motivate large numbers of people. This finding is confirmed by Figure (ref), showing the occupational degree of charisma by major occupational groups.
Interestingly, an almost monotonic pattern appears between education requirements and degree of charisma as shown in Figure (ref). The educational requirements that score the lowest on charisma is high school diploma or equivalent, whereas occupations with educational requirements of bachelor's degrees or higher are increasingly associated with charisma.
Last, we examine how charisma is connected to wage in Figure (ref) by the means of a smoothed polynomial regression. Figure (ref) indicates that there is a non-monotonic relationship between charisma and wage. Essentially, in the lowest part of the wage distribution, occupations tend to have high degrees of charisma. These occupations are, e.g., actors (presumably non-Hollywood actors), teachers, and social workers. In contrast, the occupations with the lowest score on charisma are on average found between the first and second quartile of the wage distribution, whereas occupations in the upper part of the wage distribution can be characterized by high degrees of charisma. The latter would be the occupations that require advanced degrees and leadership capabilities.
Although the term gained popularity by Goleman1995, the concept of EQ has been well-studied before starting with Maslow1950,Beldoch1964,Leuner1966. Generally speaking, EQ refers to the ability to intelligently perceive, understand, and manage emotions, and people with high EQ use this ability to guide behavior. The occupational degree of EQ is particularly interesting as EQ has been shown to correlate positively with many desirable outcomes. In summary, a recent review by Mayer2008 documents that higher EQ correlates with better social relations, better perceptions by others, better academic achievement, and better general well-being, e.g., higher life satisfaction.
We begin the analysis with Table (ref), tabulating the top and bottom ten occupations on EQ, and Figure (ref) showing EQ by occupational groups. Many top 10 occupations are either therapists, counselors, or psychologists, and these occupations truly require the ability to carry out reasoning about emotions to support others in various needs. Contrary in the bottom 10, these occupations, e.g., brickmasons, riggers, or tapers, require less emotional knowledge and more arm-hand steadiness and manual dexterity. This applies more generally to occupations within construction, maintenance, installation, and repair.
An interesting tail effect is found when considering EQ by educational requirements in Figure (ref). We generally find a monotonic relationship from low education and low EQ to high education and high EQ, but this monotonicity stops once we consider occupations that on average require a doctorate. This indicates that occupations that require doctorates rather than people with a master's degree need very specialized employees that may not have to engage too often with a larger group of coworkers rather than focusing deeply on very complex problems.
Last, we consider the usual smoothed regression against ranks in the wage distribution in Figure (ref). The EQ-wage relationship from Figure (ref) is slightly concave for low-wage occupations and slightly convex and tapering off for high-wage occupations, but generally high-wage occupations are associated with high levels of EQ. That EQ should be positively correlated with wage at the individual level is also found by Rode_2017, however, the full distribution tells us that tail effects do exists. The highly nonlinear relationship between occupational charisma and wage in Figure (ref) cannot, however, is not found EQ. This tentatively suggests that EQ is on average a more difficult characteristic to achieve compared to charisma in terms of wage returns. However, we emphasize that these relationships are mere correlations and we do not attempt to draw any causal inference.
It is inherently interesting to study occupational characteristics and how they associate with descriptive statistics of the labor market, e.g., wage, educational attainment, labor force participation rates, etc. These studies inform policy-makers on how to optimally design education and labor markets. But estimating the occupational degree of certain characteristics is challenging for many reasons. Most importantly, it may be severely difficult to validate and verify novel measures of occupational characteristics once constructed. This implicitly leaves too much room for research creativity with regards to the selection of occupational features that enter the composition of new ones.
We propose occ2vec as a fully data-driven and principal approach to quantifying occupations, which enable the measurement of any occupational characteristic of interest in a transparent and verifiable way. Using NLP, we embed more than 17,000 occupation-specific descriptors sourced from the O*NET as vectors and combine them into unique occupation vectors of high dimensions. Using an objective and reliable definition of a given target characteristic of interest, we also embed this into a high-dimensional vectors and measure the occupational degree of the given characteristic as the cosine similarity between the vectors.
Using the actual scores from O*NET on 244 attributes common to the universe of occupations as the ground truth, we extensively validate our approach by comparing our estimates to those scores. We find that our estimates explain almost 60% of the variation in the original scores and, joint with other validation exercises, we take this as evidence that our framework is capable of producing high-quality occupation vectors that accurately capture detailed aspects of occupations.
We apply occ2vec to four applications, where we study the occupational degree of two known characteristics as well as two novel. As known, or previously studied, characteristics, we consider the popular task measures (abstract, manual, and routine tasks), and exposure to AI. Our estimates of the task measures largely match the original ones. For instance, we also find that occupations involving many routine or manual tasks tend to appear in the bottom of the wage distribution. The opposite holds for abstract tasks. Regarding exposure to AI, our measure also broadly agrees with the already-proposed measures and can in fact be seen as a mix of the two most widely used ones. As novel characteristics, we consider the occupational degree of charisma and EQ. We find that occupations that score high on these attributes tend to be found within community and social service occupations, arts, entertainment, sports, and media occupations, and educational instruction occupations. High scores occur most often for occupations that require advanced degrees and in the top of the wage distribution. The most striking difference in found for occupations in the bottom of the wage distribution, where scores tend to be high on charisma but not on EQ. This suggests that wage returns to charisma and EQ could be very different.
In summary, this paper proposes a data-driven and universal approach to representing occupations as mathematical objects. This has many interesting economic applications, and particularly, we demonstrate how one may advance from purely qualitative definitions to quantitative scores at the occupational level. The occ2vec framework, therefore, opens many doors for future analyses that study how novel occupational characteristics that have previously been inaccessible to researchers matter for labor market outcomes. \phantomsection \addcontentsline{toc}{section}{References}