EconBase
← All papers

Imputation Strategies for Rightcensored Wages in Longitudinal Datasets

Jörg Drechsler, Johannes Ludsteck

arXiv 18 Feb 2025 · Econometrics · publishedJournal for Labour Market Research (2025) · 1 citations (OpenAlex)

arXiv:2502.12967 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Censoring from above is a common problem with wage information as the reported wages are typically top-coded for confidentiality reasons. In administrative databases the information is often collected only up to a pre-specified threshold, for example, the contribution limit for the social security system. While directly accounting for the censoring is possible for some analyses, the most flexible solution is to impute the values above the censoring point. This strategy offers the advantage that future users of the data no longer need to implement possibly complicated censoring estimators. However, standard cross-sectional imputation routines relying on the classical Tobit model to impute right-censored data have a high risk of introducing bias from uncongeniality (Meng, 1994) as future analyses to be conducted on the imputed data are unknown to the imputer. Furthermore, as we show using a large-scale administrative database from the German Federal Employment agency, the classical Tobit model offers a poor fit to the data. In this paper, we present some strategies to address these problems. Specifically, we use leave-one-out means as suggested by Card et al. (2013) to avoid biases from uncongeniality and rely on quantile regression or left censoring to improve the model fit. We illustrate the benefits of these modeling adjustments using the German Structure of Earnings Survey, which is (almost) unaffected by censoring and can thus serve as a testbed to evaluate the imputation procedures.

Citation extraction

47
references
66
in-text mentions
47
distinct cited
1
self-citations
9,287
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Card, D., J. Heining, and P. Kline (2013) Workplace Heterogeneity and the rise of West German Wage Inequality1.000134100%
2Dustmann, C., J. Ludsteck, and U. Schönberg (2009) Revisiting the german wage structure0.81142100%
3Meng, X.-L (1994) Multiple-Imputation Inference with Uncongenial Sources of Input (with discussion)0.64422100%
4Gartner, H (2005) The imputation of wages above the contribution limit with the german iab employment sample0.64422100%
5Chernouzhukov, V. and H. Hong (2002) Three-Step Censored Quantile Regression and Extramarital Affairs0.51121100%
6Schluter, C. and M. Trede (2023) Spatial income inequality0.51121100%
7Abowd, J., F. Kramarz, and D. Margolis (1999) High Wage Workers and High Wage Firms0.40511100%
8Anderson, A. B., A. Basilevsky, and D. P. J. Hum (1983) Handbook of Survey Research, Chapter Missing Data: A Review of the Literature, pp.\ 415–4920.40511100%
9Buchinsky, M (1994) Changes in the U.S. Wage Structure 1963 1987: Application of Quantile Regression0.40511100%
10Buchinsky, M. and J. Hahn (1998) An Alternative Estimator for the Censored Quantile Regression Model0.40511100%

Showing the top 10 of 47 scored citations.