EconBase
← All papers

CAREER: A Foundation Model for Labor Sequence Data

Keyon Vafa, Emil Palikot, Tianyu Du, Ayush Kanodia, Susan Athey, David M. Blei

arXiv 16 Feb 2022 · Machine Learning · 3 citations (OpenAlex)

arXiv:2202.08370 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Labor economists regularly analyze employment data by fitting predictive models to small, carefully constructed longitudinal survey datasets. Although machine learning methods offer promise for such problems, these survey datasets are too small to take advantage of them. In recent years large datasets of online resumes have also become available, providing data about the career trajectories of millions of individuals. However, standard econometric models cannot take advantage of their scale or incorporate them into the analysis of survey data. To this end we develop CAREER, a foundation model for job sequences. CAREER is first fit to large, passively-collected resume data and then fine-tuned to smaller, better-curated datasets for economic inferences. We fit CAREER to a dataset of 24 million job sequences from resumes, and adjust it on small longitudinal survey datasets. We find that CAREER forms accurate predictions of job sequences, outperforming econometric baselines on three widely-used economics datasets. We further find that CAREER can be used to form good predictions of other downstream variables. For example, incorporating CAREER into a wage model provides better predictions than the econometric models currently in use.

Citation extraction

66
references
135
in-text mentions
66
distinct cited
4
self-citations
6,714
main-text words

appendix boundary found by appendix_command · 50% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1J. Devlin, M. Chang, K. Lee, and K. Toutanova (2019) BERT: Pre-training of deep bidirectional transformers for language understanding1.00053100%
2Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and… (2019) Language models are unsupervised multitask learners0.92843100%
3Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever (2018) Improving language understanding by generative pre-training0.8435360%
4Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Sy… (2023) Simple and controllable music generation0.84333100%
5Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis… (2023) Starcoder: may the source be with you!0.84333100%
6Robert E Hall (1972) Turnover in the labor force0.8307457%
7Francine D Blau and Lawrence M Kahn (2017) The gender wage gap: Extent, trends, and explanations0.81142100%
8Francisco J. R. Ruiz, Susan Athey, and David M. Blei (2020) SHOPPER: A probabilistic model of consumer choice with substitutes and complements self0.7946450%
9Liangyue Li, How Jing, Hanghang Tong, Jaewon Yang, Qi He, and Bee-Ch… (2017) NEMO: Next career move prediction with contextual embedding0.7375440%
10Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jo… (2017) Attention is all you need0.7374450%

Showing the top 10 of 66 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Estimating Wage Disparities Using Foundation Models0.89474
2LABOR-LLM: Language-Based Occupational Representations with Large Language Models0.855166
3Model-Agnostic Covariate-Assisted Inference on Partially Identified Causal Effects0.40511
4Causal Inference on Outcomes Learned from Text0.40511