arXiv 12 Aug 2026 · Econometrics
arXiv:2608.11784 · PDF · Extracted main text
Outcomes are increasingly regressed on a calibrated probability vector for unobserved class membership, and that vector is often coarsened to a hard label first. Under a constant-coefficient structural mean and conditional calibration, the observed-data problem is a partially linear regression of the outcome on the probability vector; we take this reduction as the starting point and ask what coarsening costs. For any coarsening, the plug-in estimator converges to $\mathcal{A}τ$, where the coarsening operator satisfies $\mathcal{A}=I+D^{-1}\mathbb{E}[a_{h}u^{\top}]$ with $u$ the discarded signal. Coarsening is therefore free exactly when what is discarded is uncorrelated with what is kept, and is otherwise anisotropic: it distorts some contrasts far more than others. The same operator governs inference. The Wald interval built from coarsened labels has limiting coverage $Φ(z-λ)-Φ(-z-λ)$, with $λ$ the ratio of the coarsening bias to the reported standard error; because $\mathcal{A}$ and that standard error depend on observables alone, the coverage implied by the estimated index can be approximated before the interval is reported. Simulations show severe coverage loss after argmax coarsening, and three real-data audits exhibit the direction-specific distortion that hard labels induce.
appendix boundary found by appendix_titled_section at “Supplementary Materials” · 51% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Chen, J., Kallus, N., Mao, X., Svacha, G., and Udell, M (2019) Fairness under unawareness: Assessing disparity when protected class is unobserved | 1.000 | 5 | 3 | 100% |
| 2 | Kallus, N., Mao, X., and Zhou, A (2022) Assessing algorithmic fairness with unobserved protected class using data combination | 0.928 | 4 | 3 | 100% |
| 3 | Kurbucz, M. T (2026) When to trust confidence thresholding: Calibration diagnostics for pseudo-labelled regression self | 0.928 | 4 | 3 | 100% |
| 4 | Armstrong, T. B. and Kolesár, M (2021) Sensitivity analysis using approximate moment condition models | 0.843 | 3 | 3 | 100% |
| 5 | Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters | 0.737 | 3 | 3 | 67% |
| 6 | Angelopoulos, A. N., Bates, S., Fannjiang, C., Jordan, M. I., and Zr… (2023) Prediction-powered inference | 0.737 | 3 | 2 | 100% |
| 7 | Robinson, P. M (1988) Root-N-consistent semiparametric regression | 0.737 | 3 | 2 | 100% |
| 8 | Battaglia, L., Christensen, T., Hansen, S., and Sacher, S (2024) Inference for regression with variables generated by AI or machine learning | 0.644 | 2 | 2 | 100% |
| 9 | Lee, D.-H (2013) Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks | 0.644 | 2 | 2 | 100% |
| 10 | Sohn, K., Berthelot, D., Li, C.-L., Zhang, Z., Carlini, N., Cubuk, E… (2020) FixMatch: Simplifying semi-supervised learning with consistency and confidence | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 33 scored citations.