EconBase
← All papers

Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss

Marcell T. Kurbucz

arXiv 12 Aug 2026 · Econometrics

arXiv:2608.11784 · PDF · Extracted main text

Abstract

Outcomes are increasingly regressed on a calibrated probability vector for unobserved class membership, and that vector is often coarsened to a hard label first. Under a constant-coefficient structural mean and conditional calibration, the observed-data problem is a partially linear regression of the outcome on the probability vector; we take this reduction as the starting point and ask what coarsening costs. For any coarsening, the plug-in estimator converges to $\mathcal{A}τ$, where the coarsening operator satisfies $\mathcal{A}=I+D^{-1}\mathbb{E}[a_{h}u^{\top}]$ with $u$ the discarded signal. Coarsening is therefore free exactly when what is discarded is uncorrelated with what is kept, and is otherwise anisotropic: it distorts some contrasts far more than others. The same operator governs inference. The Wald interval built from coarsened labels has limiting coverage $Φ(z-λ)-Φ(-z-λ)$, with $λ$ the ratio of the coarsening bias to the reported standard error; because $\mathcal{A}$ and that standard error depend on observables alone, the coverage implied by the estimated index can be approximated before the interval is reported. Simulations show severe coverage loss after argmax coarsening, and three real-data audits exhibit the direction-specific distortion that hard labels induce.

Citation extraction

33
references
58
in-text mentions
33
distinct cited
2
self-citations
9,438
main-text words

appendix boundary found by appendix_titled_section at “Supplementary Materials” · 51% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chen, J., Kallus, N., Mao, X., Svacha, G., and Udell, M (2019) Fairness under unawareness: Assessing disparity when protected class is unobserved1.00053100%
2Kallus, N., Mao, X., and Zhou, A (2022) Assessing algorithmic fairness with unobserved protected class using data combination0.92843100%
3Kurbucz, M. T (2026) When to trust confidence thresholding: Calibration diagnostics for pseudo-labelled regression self0.92843100%
4Armstrong, T. B. and Kolesár, M (2021) Sensitivity analysis using approximate moment condition models0.84333100%
5Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters0.7373367%
6Angelopoulos, A. N., Bates, S., Fannjiang, C., Jordan, M. I., and Zr… (2023) Prediction-powered inference0.73732100%
7Robinson, P. M (1988) Root-N-consistent semiparametric regression0.73732100%
8Battaglia, L., Christensen, T., Hansen, S., and Sacher, S (2024) Inference for regression with variables generated by AI or machine learning0.64422100%
9Lee, D.-H (2013) Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks0.64422100%
10Sohn, K., Berthelot, D., Li, C.-L., Zhang, Z., Carlini, N., Cubuk, E… (2020) FixMatch: Simplifying semi-supervised learning with consistency and confidence0.64422100%

Showing the top 10 of 33 scored citations.