EconBase
← Back to paper

Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

61,300 characters

Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss




\def\spacingset#1{\renewcommand{\baselinestretch}
{#1}\small\normalsize} \spacingset{1}

\if10
{
  \title{\bf Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss}
  \author{Author(s)\\ Department, University}
  \maketitle
} \fi

\if00
{
  \bigskip\bigskip\bigskip
  \begin{center}
    {\LARGE\bf Coarsening Latent-Class Probabilities:\\[0.5ex] Directional Distortion and Coverage Loss}
  \end{center}
  \medskip
  \begin{center}
    Marcell T.\ Kurbucz\\
    Institute for Global Prosperity, The Bartlett,\\
    University College London\\
    \texttt{[email removed]}\\[0.5ex]
    \today
  \end{center}
  \medskip
} \fi

\bigskip
\begin{abstract}
\noindent Outcomes are increasingly regressed on a calibrated probability vector for unobserved class membership, and that vector is often coarsened to a hard label first. Under a constant-coefficient structural mean and conditional calibration, the observed-data problem is a partially linear regression of the outcome on the probability vector; we take this reduction as the starting point and ask what coarsening costs. For any coarsening, the plug-in estimator converges to $\mathcal{A}\tau$, where the coarsening operator satisfies $\mathcal{A}=I+D^{-1}\mathbb{E}[a_{h}u^{\!\top}]$ with $u$ the discarded signal. Coarsening is therefore free exactly when what is discarded is uncorrelated with what is kept, and is otherwise anisotropic: it distorts some contrasts far more than others. The same operator governs inference. The Wald interval built from coarsened labels has limiting coverage $\Phi(z-\lambda)-\Phi(-z-\lambda)$, with $\lambda$ the ratio of the coarsening bias to the reported standard error; because $\mathcal{A}$ and that standard error depend on observables alone, the coverage implied by the estimated index can be approximated before the interval is reported. Simulations show severe coverage loss after argmax coarsening, and three real-data audits exhibit the direction-specific distortion that hard labels induce.
\end{abstract}

\noindent
{\it Keywords:} classify-analyze; double machine learning; measurement error; partial identification; proxy variables; pseudo-labeling
\vfill

\newpage
\spacingset{1.15}

\section{Introduction}
\label{sec:intro}

A recurring problem in empirical work is to measure how membership in a latent group affects an outcome when the group indicator is never observed but a \emph{calibrated probability} for it is. Such probabilities come from a trained classifier or a model-based score: an unobserved protected attribute proxied for a fairness audit \citep{kallus2022,chen2019}, disease status from an imaging classifier, poverty status from administrative predictors. The binary version of this problem admits a closed-form treatment: when a scalar calibrated score $p$ satisfies the conditional calibration condition $\mathbb{E}[G\mid p,X]=p$, the structural effect $\tau$ of the latent indicator $G\in\{0,1\}$ is point-identified by a moment equation whose denominator is a residual-score variance. Identification fails exactly when that variance vanishes. This is the $K=1$ special case of the results developed below (see \citealp{kurbucz2026binary}, for the binary theory, and \citealp{kurbucz2026kbs}, for its confidence-thresholding diagnostics).

Most applications, however, are not binary. Fairness audits compare several protected groups; prevalence studies distinguish multiple disease subtypes; classifiers emit a probability vector over $K+1$ mutually exclusive classes. The natural object is then a vector $\tau\in\mathbb{R}^{K}$ of group effects and a vector-valued calibrated score $p$ on the probability simplex. This paper asks: \emph{what does coarsening that vector cost the coefficient and the interval reported from it, along which directions, and can the cost be known before either is reported?}

The existing answers repair rather than diagnose. One line of work documents the damage to an aggregate quantity, where a hard label demonstrably distorts a count or a group mean \citep{dong2025,chen2019}; another corrects the downstream estimate inside a latent-class model, using an estimated matrix of classification errors \citep{bolck2004}; a third buys validity back with a gold-standard subsample \citep{angelopoulos2023}. Each presupposes either a model the analyst has not fitted or data the analyst does not have, and none of them answers the question that arises once an interval is already on the page: what is that interval worth?

The argument rests on a constant-coefficient structural mean and on conditional calibration. Together they imply $\mathbb{E}[Y\mid p,X]=\mu(X)+\tau^{\!\top} p$, so on observables the problem is a partially linear regression of the outcome on $p$ \citep{robinson1988}. We state this at the outset because it is what the assumptions buy: conditional calibration is precisely the restriction that turns unobserved membership into an observable regression, without instruments, repeated measurements or a validation sample.

Two objects organize the answer, and the first is a matrix. Define the score residual $a=p-r(X)$ with $r(X)=\mathbb{E}[p\mid X]$, and the \emph{residual-score covariance matrix}
\begin{equation}
  M \;=\; \mathbb{E}\!\left[(p-r(X))(p-r(X))^{\!\top}\right] \;=\; \mathbb{E}[\operatorname{Cov}(p\mid X)]
  \;\in\;\mathbb{R}^{K\times K}.
  \label{eq:M-intro}
\end{equation}
When $K=1$ this recovers the scalar residual-score variance $V^{*}=\mathbb{E}[(p-r(X))^2]$ of the binary case.

The paper is about what happens when that regression is not the one an analyst runs. In routine practice the probability vector is coarsened before use: replaced by an argmax label \citep{lee2013}, by a confidence-thresholded label \citep{sohn2020}, or by a rounded summary. The second object is the operator that this substitution induces, and our first result identifies it: for \emph{any} such coarsening $h$, the plug-in estimator converges to $\mathcal{A}_h\tau$ with
\begin{equation}
  \mathcal{A}_h \;=\; I + D_h^{-1}\mathbb{E}[a_h u^{\!\top}],
  \qquad a_h=h-\mathbb{E}[h\mid X],\quad u=a-a_h.
  \label{eq:A-intro}
\end{equation}
The distortion is the covariance between the retained signal and the discarded one: coarsening costs nothing exactly when the two are uncorrelated, and otherwise acts unevenly across directions, so the damage to one contrast says little about the damage to another. This is the multivariate phenomenon the binary case cannot express, where the same object is a scalar dilution factor \citep{kurbucz2026kbs}.

The second result concerns the interval rather than the point. An analyst who coarsens and then forms the usual sandwich reports a Wald interval whose limiting coverage is $\Phi(z-\lambda_v)-\Phi(-z-\lambda_v)$, with $\lambda_v$ the ratio of the coarsening bias along $v$ to the reported standard error (Proposition~\ref{prop:coverage}). Two features make this usable rather than merely cautionary. Coverage falls as the coarsened estimator becomes more precise, because $\lambda_v$ grows as the reported standard error shrinks; and both ingredients are observable, since $\mathcal{A}_h$ is identified from $(p,X)$ alone and the reported standard error from $(Y,p,X)$. One can therefore estimate what the interval will cover before reporting it, with no class labels. Reading $\mathcal{A}_h\tau$ as a distortion of the latent effect does require the identifying model, and the index below also requires an estimate of $\tau$. In the leading design an argmax interval is $39\%$ narrower than the correct one and covers $1\%$ of the time.

Two further results tie $\mathcal{A}_h$ to $\tau$ itself rather than to whatever regression coefficient sits in its place. Along an eigen-direction with vanishing eigenvalue the moment restriction leaves $\tau$ set-valued, and the failure is observational rather than an artifact of one moment (Theorem~\ref{thm:spectral}). When calibration itself fails within a budget $\delta$, the induced bias along $v_j$ is bounded by $\delta\|\tau\|_1\lambda_j^{-1/2}$ (Proposition~\ref{prop:cal}), following the approximate-moment-condition approach of \citet{armstrong2021}. The same spectrum that grades precision grades fragility to miscalibration; coarsening is direction-specific as well, though not ordered by those eigenvalues.

In the proxy-based disparity audit, the setting nearest to ours, \citet{chen2019} contrast hard thresholded imputation with a probability-weighted estimator of a binary demographic disparity. \citet{kallus2022} show that when the outcome and the protected class are never observed together, the coupling between them within proxy cells is unrestricted and common disparity measures are only partially identified. Our exclusion restriction constrains exactly that coupling, which is what turns partial into point identification; along a collapsed eigen-direction we return to a partially identified regime.

On North Carolina voter records, \citet{dong2025} attribute $11.9$ of a $28.2$-point undercount of African-American voters in a commercial voter file to argmax labeling rather than to miscalibration, and \citet{xin2026} obtain, to first order, coefficients that mix the true group effects through a confusion matrix. The estimands vary---a group-mean disparity, an aggregate count, and in the last case a regression coefficient as here---but the operator below is exact rather than first-order, covers any coarsening rather than a misclassified category, and is identified without ground-truth labels. None of this work asks whether the reported interval remains valid.

The misclassification literature identifies regression effects of a mismeasured binary regressor \citep{mahajan2006} and shows that misclassification attenuates them \citep{lewbel2007}; \citet{molinari2008} bounds the distribution of the underlying variable through the matrix of misclassification probabilities. An older antecedent sits in the latent-class literature, where the classify-analyze strategy classifies units from their posterior class probabilities---most simply by taking the class with the largest one---uses the assignment downstream as though it were the truth, and finds the association attenuated. The corrections either invert the classification-error matrix \citep{bolck2004,vermunt2010} or enlarge the classification model \citep{bray2015}. All of this presupposes a corrupted label or a fitted latent-class model. We observe a calibrated probability vector, so the coarsening is chosen by the analyst rather than imposed by the data. That is what makes $\mathcal{A}_h$ estimable with no first-stage model, and it moves the question from correcting the estimate to assessing the interval.

Recovering the latent structure itself calls for devices unavailable here: diagonalizing integral operators under injectivity conditions \citep{hu2008,schennach2016}, or many conditionally independent measurements \citep{allman2009,kasahara2009}. The data here carry a single probability vector, and the analysis rests instead on the partially linear model \citep{robinson1988} with double machine learning \citep{chernozhukov2018}. Recent work on regressors produced by machine learning \citep{battaglia2024} and on inference that leans on a gold-labeled subsample \citep{angelopoulos2023} shares the concern of the coverage result here. In both, however, the distortion originates in estimation error, whereas a deterministic coarsening of an exactly observed vector calls for a diagnostic rather than a correction.

Section~\ref{sec:model} sets up the model and records the reduction to a partially linear regression. Section~\ref{sec:ident} states what that regression identifies and where it fails; Section~\ref{sec:inference} gives orthogonality, the sandwich limit, and the efficiency bound. Section~\ref{sec:atten} introduces the coarsening operator and Section~\ref{sec:coverage} the coverage of the resulting interval. Section~\ref{sec:sim} reports simulations, Section~\ref{sec:adult} a UCI Adult disparity audit, and Section~\ref{sec:recal} a sensitivity bound for miscalibration together with a labeled-subsample experiment separating repairable calibration failure from an exclusion violation. Section~\ref{sec:bisg} exhibits a surname-based voter-turnout audit ($n=464{,}700$). Section~\ref{sec:disc} discusses scope. Proofs, additional experiments, and a land-cover audit are in the supplement.

\section{Model and Assumptions}
\label{sec:model}

We observe i.i.d.\ draws $(Y_i,X_i,p_i)\in\mathbb{R}\times\mathcal{X}\times\Delta_K$, where $\Delta_K=\{p\in[0,1]^{K}:\mathbf 1^{\!\top} p\le 1\}$ collects the first $K$ coordinates of a $(K{+}1)$-class probability vector; the omitted reference probability is $p_{K+1}=1-\mathbf 1^{\!\top} p$. On the same space there is an unobserved one-hot membership vector $g\in\{0,1\}^{K}$, $g_k=\mathbf{1}\{G=k\}$, $G\in\{1,\dots,K+1\}$, with $K+1$ the reference class. Write $m(x)=\mathbb{E}[Y\mid X=x]$, $r(x)=\mathbb{E}[p\mid X=x]\in\mathbb{R}^{K}$, and define
\begin{equation}
  a \;=\; p-r(X)\in\mathbb{R}^{K}, \qquad R \;=\; Y-m(X),
  \label{eq:resid}
\end{equation}
so $\mathbb{E}[a\mid X]=0$ and $\mathbb{E}[R\mid X]=0$. The residual-score covariance matrix $M=\mathbb{E}[a a^{\!\top}]$ of \eqref{eq:M-intro} is symmetric positive semidefinite.

\begin{assumption}[Structural conditional mean]
\label{ass:mean}
There exist a measurable $\mu:\mathcal{X}\to\mathbb{R}$ and a vector $\tau\in\mathbb{R}^{K}$ with $\mathbb{E}[Y\mid G,p,X]=\mu(X)+\tau^{\!\top} g$ almost surely; here $\tau_k$ is the conditional mean contrast between class $k$ and the reference class $K{+}1$, holding $X$ fixed.
\end{assumption}

\begin{assumption}[Conditional calibration]
\label{ass:cal}
$\mathbb{E}[g\mid p,X]=p$ almost surely.
\end{assumption}

\begin{assumption}[Non-degeneracy]
\label{ass:rank}
$M=\mathbb{E}[(p-r(X))(p-r(X))^{\!\top}]$ is nonsingular ($\lambda_{\min}(M)>0$).
\end{assumption}

\begin{assumption}[Moments]
\label{ass:mom}
$\mathbb{E}[Y^4]<\infty$ (since $p\in\Delta_K$ is bounded, all moments of $p$ are finite).
\end{assumption}

Assumption~\ref{ass:mean} has two parts: the latent-class effects are constant in $X$, and the probability vector is mean-independent of $Y$ given $(G,X)$. The latter is an \emph{exclusion restriction}. The features used to build $p$ may predict membership, but conditional on $G$ and the downstream controls $X$ they carry no further mean-relevant information about $Y$. We use ``effect'' for this structural conditional-mean contrast, not a causal effect absent further assumptions; the supplement treats the marginal contrast, which is also point-identified and differs from $\tau_k$ by an explicit compositional term, and the case of effects varying with $X$. Assumption~\ref{ass:cal} is the sole link between the unobserved $g$ and the observed $p$, and is stronger than ordinary multiclass calibration: it requires the probability vector to be unbiased for membership after conditioning on $X$, the analogue of an identifying exogeneity condition. Sub-population miscalibration in $X$ therefore violates it, and Proposition~\ref{prop:cal} shows the resulting bias is amplified by small eigenvalues of $M$. Assumption~\ref{ass:rank} is the multivariate non-degeneracy condition.

\begin{lemma}[Residual decomposition]
\label{lem:resid}
Under Assumptions~\ref{ass:mean}--\ref{ass:mom},
\begin{equation}
  R \;=\; \tau^{\!\top}\!\left(g-r(X)\right) + \varepsilon,
  \qquad \mathbb{E}[\varepsilon\mid G,p,X]=0.
  \label{eq:lem-resid}
\end{equation}
\end{lemma}

Lemma~\ref{lem:resid} (proved in the supplement) is the structural backbone: the entire predictable part of the outcome residual is driven by the deviation $g-r(X)$ of true membership from its probability-implied expectation.

\subsection{What the two assumptions buy, and what they do not}
\label{sec:reduction}
It is worth stating at the outset what Assumptions~\ref{ass:mean} and~\ref{ass:cal} accomplish at the level of the observed data. Taking $\mathbb{E}[\,\cdot\mid p,X\,]$ in Assumption~\ref{ass:mean} and substituting Assumption~\ref{ass:cal},
\begin{equation}
  \mathbb{E}[Y\mid p,X] \;=\; \mu(X)+\tau^{\!\top}\mathbb{E}[g\mid p,X] \;=\; \mu(X)+\tau^{\!\top} p.
  \label{eq:plr}
\end{equation}
The latent class has disappeared: on observables the model is a partially linear regression of $Y$ on the probability vector $p$ with nonparametric component $\mu$ \citep{robinson1988}. Everything that follows about point identification, orthogonality and the sandwich limit is the corresponding statement for that regression, with $M$ the covariance matrix of the partialed-out regressor and $\tau$ the partial regression coefficient.

We record \eqref{eq:plr} explicitly because it is the content of the two assumptions, not an obstacle to it. It also fixes the scope of the paper. Sections~\ref{sec:ident}--\ref{sec:inference} record what \eqref{eq:plr} implies, in the form needed later; the contribution begins in Section~\ref{sec:atten}, where the probability vector is replaced by a coarsened version of itself and \eqref{eq:plr} no longer applies.

\section{Identification and its Spectral Geometry}
\label{sec:ident}

\subsection{The matrix moment identity}

\begin{theorem}[Identification]
\label{thm:ident}
Under Assumptions~\ref{ass:mean}, \ref{ass:cal} and \ref{ass:mom},
\begin{equation}
  \mathbb{E}[a R] \;=\; M\,\tau.
  \label{eq:moment}
\end{equation}
If in addition Assumption~\ref{ass:rank} holds, then $\tau=M^{-1}\mathbb{E}[aR]$ is point-identified from the joint law of $(Y,X,p)$.
\end{theorem}

The proof is in the supplement. Equation~\eqref{eq:moment} is the matrix analogue of the scalar identity $\mathbb{E}[(p-r(X))R]=V^{*}\tau$ (the $K=1$ case, with $V^{*}=\mathbb{E}[(p-r(X))^2]$): the scalar residual-score variance $V^{*}$ becomes the covariance matrix $M$, and the scalar covariance $\mathbb{E}[(p-r)R]$ becomes the cross-moment vector $\mathbb{E}[aR]$. The estimand is a multivariate partial regression of $R$ on the residualized score $a$.

\subsection{Spectral characterization of identification failure}

Let $M=\sum_{j=1}^{K}\lambda_j v_j v_j^{\!\top}$ be the spectral decomposition, with $0\le\lambda_1\le\cdots\le\lambda_K$ and orthonormal eigenvectors $v_j$.

\begin{theorem}[Spectral identification]
\label{thm:spectral}
Under Assumptions~\ref{ass:mean}, \ref{ass:cal} and \ref{ass:mom}:
\begin{enumerate}[(a)]
\item The moment equation \eqref{eq:moment} identifies $M\tau$, hence the
orthogonal projection of $\tau$ onto $\mathcal{R}(M)=\mathrm{span}\{v_j:\lambda_j>0\}$.
The full vector $\tau$ is point-identified if and only if $\lambda_1>0$.
\item If $\operatorname{rank}(M)=K-q<K$, the moment \eqref{eq:moment} does not
point-identify $\tau$: it determines only the projection
$P_{\mathcal{R}(M)}\tau=M^{+}\mathbb{E}[aR]$ (with $M^{+}$ the Moore--Penrose
inverse), and its solution set is the affine subspace $M^{+}\mathbb{E}[aR]+\mathcal{N}(M)$.
The components of $\tau$ in $\mathcal{N}(M)$ are unrestricted by the moment.
\item $\lambda_j=0$ if and only if $v_j^{\!\top} p=v_j^{\!\top} r(X)$ almost surely; that is,
the score combination $v_j^{\!\top} p$ is a deterministic function of $X$.
\item The failure in (b) is genuinely observational, not an artifact of one
moment equation. Let $\tilde m(p,X)=\mathbb{E}[Y\mid p,X]$, $u=Y-\tilde m(p,X)$,
$g_0(u)=u/(1+|u|)$, and suppose there is $\underline\kappa>0$ with
$\kappa(p,X):=\mathbb{E}[u\,g_0(u)\mid p,X]\geq\underline\kappa$ a.s.\ and
$\|\tau\|_\infty\leq\underline\kappa/8$. Then for every
$z\in\mathcal{N}(M)$ with $\|z\|_\infty=1$ and every $c$ with
$|c|\leq\bar c:=\underline\kappa/8$, there exists a model satisfying
Assumptions~\ref{ass:mean}, \ref{ass:cal} and~\ref{ass:mom} with coefficient
vector $\tau+cz$ in which the observables $(Y,X,p)$ have exactly their
original joint distribution. The identified set is convex, contains the
segment $\{\tau+cz:|c|\leq\bar c\}$ along every null direction, and---whenever
each class probability is bounded away from zero on a positive-probability
set on which $\mathbb{E}[u^2\mid p,X]$ is essentially bounded---it is bounded, hence
a strict subset of the affine subspace in (b).
\end{enumerate}
\end{theorem}

Theorem~\ref{thm:spectral} (proved in the supplement) records where the regression \eqref{eq:plr} stops identifying $\tau$. Part~(c) says a direction collapses precisely when the corresponding linear combination of probabilities carries no residual variation beyond $X$. Part~(b) says the failure is local: the effect remains point-identified along every other direction. Part~(d) upgrades the moment-level statement to observational equivalence, by tilting the conditional class probabilities with a bounded, conditionally mean-zero transform of the outcome residual. The construction and the resulting bounded segment are given in the supplement, together with the reasons the conditions are sufficient rather than sharp. When $K=1$ identification is all or nothing; with several classes an analyst can have sharp inference on some contrasts and none on others in the same study, and the spectrum of $\widehatM$ says which.

In the disparity-audit literature, part~(d) is the spectrum-indexed counterpart of the sharp partial-identification sets of \citet{kallus2022}, the difference being that Assumption~\ref{ass:mean} restricts the outcome--membership coupling throughout, so that only the collapsed directions remain set-identified.

\begin{remark}[Clustered eigenvalues]
\label{rem:cluster}
When eigenvalues cluster, the sample eigenvectors of $\widehatM$ are unstable, and direction-specific quantities should be reported for the spanned eigenspace rather than a single eigenvector; in the simulations, inference on a fixed direction is unaffected even when two small eigenvalues nearly coincide.
\end{remark}

\section{Estimation and Inference}
\label{sec:inference}

\subsection{Estimator and Neyman orthogonality}

With nuisances $m,r$ estimated by cross-fitting, the moment is
\begin{equation}
  \psi(W;\tau,m,r)\;=\;a\,(R-a^{\!\top}\tau),\qquad a=p-r(X),\ R=Y-m(X),
  \label{eq:psi}
\end{equation}
and the estimator solves $n^{-1}\sum_i\hat a_i(\hat R_i-\hat a_i^{\!\top}\hat\tau)=0$, i.e.
\begin{equation}
  \hat\tau \;=\; \Bigl(n^{-1}\textstyle\sum_i \hat a_i\hat a_i^{\!\top}\Bigr)^{-1}
                 \Bigl(n^{-1}\textstyle\sum_i \hat a_i\hat R_i\Bigr).
  \label{eq:est}
\end{equation}
We treat the score $p$ as an externally supplied calibrated output; the only estimated nuisances are $m$ and $r$. When $p$ is itself produced by a classifier trained in-sample, cross-fitting the classifier (as in Section~\ref{sec:adult}) controls overfitting, but our asymptotics condition on the score-generating mechanism.

\begin{proposition}[Full Neyman orthogonality]
\label{prop:ortho}
Under Assumptions~\ref{ass:mean}--\ref{ass:mom}, the moment \eqref{eq:psi} is Neyman-orthogonal with respect to both nuisances: the Gateaux derivatives of $\mathbb{E}[\psi]$ in the $m$- and $r$-directions vanish at the truth.
\end{proposition}

Proposition~\ref{prop:ortho} (proved in the supplement) is what makes \eqref{eq:est} directly compatible with machine-learning nuisance estimators at the usual product rate \citep{chernozhukov2018}: the symmetric residual form \eqref{eq:psi} is orthogonal in both $m$ and $r$.

\subsection{Multivariate sandwich CLT}

\begin{theorem}[Multivariate CLT]
\label{thm:clt}
Under Assumptions~\ref{ass:mean}--\ref{ass:mom}, i.i.d.\ sampling, and nuisance estimates satisfying $\|\hat m-m\|_{2}\,\|\hat r-r\|_{2}+\|\hat r-r\|_{2}^{2}=o_P(n^{-1/2})$ with cross-fitting,
\begin{equation}
  \sqrt{n}\,(\hat\tau-\tau)\;\xrightarrow{\;d\;}\;\mathcal{N}\!\left(0,\;\Omega\right),
  \qquad \Omega=M^{-1}\Sigma\,M^{-1},
  \quad \Sigma=\mathbb{E}\!\left[u^2\,a a^{\!\top}\right],\ u=R-a^{\!\top}\tau.
  \label{eq:clt}
\end{equation}
Write $\widehat\Sigma=n^{-1}\sum_i\hat u_i^2\,\hat a_i\hat a_i^{\!\top}$ with $\hat u_i=\hat R_i-\hat a_i^{\!\top}\hat\tau$, $\widehat\Omega=\widehatM^{-1}\widehat\Sigma\,\widehatM^{-1}$, and the finite-sample variance estimate $\widehat V=\widehat\Omega/n$. Then $\widehat\Omega\xrightarrow{\;p\;}\Omega$, Wald intervals $\hat\tau_k\pm z_{1-\alpha/2}\sqrt{\widehat V_{kk}}$ and the ellipsoid $(\hat\tau-\tau)^{\!\top}\widehat V^{-1}(\hat\tau-\tau)\le\chi^2_{K,1-\alpha}$ have asymptotic coverage $1-\alpha$. Writing $\sigma^2_j=v_j^{\!\top}\Sigma v_j$, the asymptotic variance along eigen-direction $j$ is $v_j^{\!\top} \Omega v_j=\sigma_j^2/\lambda_j^2$, so precision degrades as $\lambda_j\to0$, quantifying Theorem~\ref{thm:spectral}.
\end{theorem}

The proof (in the supplement) is a standard $Z$-estimator argument using the orthogonality of Proposition~\ref{prop:ortho} to absorb the nuisance estimation error. The directional variance formula $v_j^{\!\top} \Omega v_j=\sigma_j^2/\lambda_j^2$ is the inferential face of the spectral geometry: a weakly-identified direction ($\lambda_j$ small) yields a wide confidence interval, while well-identified directions remain sharp. The exponent depends on how the moment noise behaves: with locally homoskedastic noise $\Sigma\approx\sigma^2M$ giving $v_j^{\!\top}\Omega v_j\approx\sigma^2/\lambda_j$, while at fixed $\Sigma$ it is $\lambda_j^{-2}$; the experiments show the realized scaling sits between $\lambda_j^{-1}$ and $\lambda_j^{-2}$ because $\Sigma$ co-varies with $M$ across designs (Section~\ref{sec:sim}). A companion result in the supplement gives the direction-specific rate $\sqrt{n\lambda_{v,n}}$ at which a contrast is learned as its eigenvalue drifts to zero, and shows that the Wald interval remains correctly calibrated along such a direction under conditional homoskedasticity.

\subsection{Semiparametric efficiency}
\label{sec:eff}
\noindent By Theorem~\ref{thm:ident}, $\tau$ solves the conditional moment restriction $\mathbb{E}[\,R-a^{\!\top}\tau\mid p,X\,]=0$, and the efficiency theory for such restrictions \citep{chamberlain1987} pins down the best achievable variance.

\begin{theorem}[Efficiency bound and optimal weighting]
\label{thm:eff}
Let $\sigma^{2}(p,X)=\operatorname{Var}(R\mid p,X)$, let $\tilde p(X)=\mathbb{E}[\sigma^{-2}(p,X)\,p\mid X]/\mathbb{E}[\sigma^{-2}(p,X)\mid X]$ be the precision-weighted conditional mean of $p$, and set $\tilde a=p-\tilde p(X)$. Under Assumptions~\ref{ass:mean}--\ref{ass:mom}, the semiparametric efficiency bound for $\tau$ is
\begin{equation}
  V_{\mathrm{eff}}=\bigl(\mathbb{E}[\sigma^{-2}(p,X)\,\tilde a\tilde a^{\!\top}]\bigr)^{-1},
  \label{eq:veff}
\end{equation}
attained by the estimator solving $\sum_i \sigma^{-2}(p_i,X_i)\,\tilde a_i\,(R_i-a_i^{\!\top}\tau)=0$, with $\sigma^2$, $\tilde p$, $m$ and $r$ replaced by cross-fitted estimates. The unweighted estimator \eqref{eq:est} attains \eqref{eq:veff} if and only if $\sigma^{2}(p,X)\,a=C\tilde a$ almost surely for some nonsingular constant matrix $C$; a constant $\sigma^{2}$ is the leading sufficient case. In general $V=M^{-1}\Sigma\,M^{-1}\succeq V_{\mathrm{eff}}$, with strict inequality along some direction under generic heteroskedasticity.
\end{theorem}

The proof (in the supplement) is a Chamberlain calculation carried out under the constraint that $\mu$ is unknown. An instrument $h(p,X)$ leaves the moment $\mathbb{E}[h\,(Y-\mu(X)-\tau^{\!\top} p)]$ insensitive to perturbations of $\mu$ only if $\mathbb{E}[h\mid X]=0$, and since $\sigma^{2}$ depends on $p$, the naive choice $h=\sigma^{-2}a$ violates that condition; recentering $p$ at $\tilde p(X)$ restores it, and $\sigma^{-2}\tilde a$ is optimal among the instruments that satisfy it. Heteroskedasticity is intrinsic here: since $u=\tau^{\!\top}(g-p)+\varepsilon$ and $\mathbb{E}[\varepsilon\mid G,p,X]=0$ kills the cross term, $\sigma^{2}(p,X)=\tau^{\!\top}(\operatorname{diag}(p)-pp^{\!\top})\tau+\operatorname{Var}(\varepsilon\mid p,X)$, so the latent-membership term varies with $p$ even when $\operatorname{Var}(\varepsilon\mid p,X)$ is constant, and $\tilde p$ then departs from $r$. Efficiency therefore costs two nuisances beyond $(m,r)$, the conditional variance and the precision-weighted conditional mean. The unweighted estimator, used in the rest of the paper, needs neither and is orthogonal by Proposition~\ref{prop:ortho}.

\section{Coarsening the Probability Vector}
\label{sec:atten}

Equation~\eqref{eq:plr} describes the regression the analyst \emph{could} run. It is common to run a different one. A pervasive practice when probability vectors feed a downstream analysis is to discard them and keep a coarser object: an argmax one-hot label \citep{lee2013}, retained in some variants only when its predicted probability clears a confidence threshold \citep{sohn2020}, a top-$k$ summary, or a rounded version of $p$. This section characterizes what that substitution costs.

\begin{definition}
A \emph{coarsening} is a measurable map $h=h(p,X)$ taking values in the same coordinates as membership, $h:\Delta_K\times\mathcal{X}\to\Delta_K$ or into the one-hot vertices. Write $a_h=h-\mathbb{E}[h\mid X]$ and let $\hat\tau_h$ solve $n^{-1}\sum_i\hat a_{h,i}(\hat R_i-\hat a_{h,i}^{\!\top}\tau)=0$.
\end{definition}

The requirement that $h$ live in the coordinates of $g$ is what makes the comparison with $\tau$ meaningful: a linear change of basis $h=Bp$ would return coefficients in transformed units rather than a distorted version of the same object. Coarsening in this sense is the operation studied under a different guise in the coarse-data literature \citep{heitjan1991}; here the coarsening is deliberate and its target is a downstream coefficient.

\begin{theorem}[Coarsening operator]
\label{thm:atten}
Under Assumptions~\ref{ass:mean}--\ref{ass:mom}, for any coarsening $h$ with $D_h=\mathbb{E}[a_ha_h^{\!\top}]$ nonsingular,
\begin{equation}
  \operatorname{plim}\,\hat\tau_{h} \;=\; \mathcal{A}_h\,\tau,
  \qquad
  \mathcal{A}_h \;=\; D_h^{-1}C_h,\quad
  C_h=\mathbb{E}[a_h a^{\!\top}].
  \label{eq:atten}
\end{equation}
\end{theorem}

The proof (in the supplement) is two steps: $\mathbb{E}[a_h\varepsilon]=0$ because $a_h$ is $\sigma(p,X)$-measurable, and conditioning on $(p,X)$ turns $\mathbb{E}[a_h(g-r(X))^{\!\top}]$ into $\mathbb{E}[a_ha^{\!\top}]$ by Assumption~\ref{ass:cal}. Nonsingularity of $D_h$ requires the coarsened label to retain residual variation given $X$; it fails, for instance, if the argmax never selects some class.

Writing $u=a-a_h$ for the signal the coarsening discards gives the form we use throughout.

\begin{corollary}[Distortion is covariance with the discarded signal]
\label{cor:decomp}
$\mathcal{A}_h = I + D_h^{-1}\mathbb{E}[a_h u^{\!\top}]$. Hence $\mathcal{A}_h=I$ if and only if $\mathbb{E}[a_hu^{\!\top}]=0$, and $\mathcal{A}_h=c\,I$ for a scalar $c$ if and only if $\mathbb{E}[a_hu^{\!\top}]=(c-1)D_h$.
\end{corollary}

Coarsening is free exactly when what it throws away is uncorrelated with what it keeps. That condition is restrictive, and the second part of the corollary says that even proportional shrinkage is non-generic: proportional shrinkage requires a restrictive covariance condition that the designs and audits below never satisfy.

\begin{corollary}[Directional retention]
\label{cor:aniso}
Let $s_j=v_j^{\!\top} \mathcal{A}_h v_j$ be the retention coefficient along the $j$th eigen-direction of $M$. Coarsening acts uniformly on $\tau$ only when $\mathcal{A}_h=cI$, equivalently when $\mathbb{E}[a_hu^{\!\top}]=(c-1)D_h$; outside that case it is direction-specific, and the $s_j$ need neither coincide across $j$ nor lie in $[0,1]$.
\end{corollary}

Coarsening therefore need not act on the effect vector uniformly. In every design and audit we examine, the $s_j$ of a non-trivial coarsening lie in $(0,1)$ and differ across directions, so that coarsening costs some contrasts much more than others; the corollary does not assert this in general, and the diagnostic below reports the $s_j$ rather than presuming their range or their ordering.

Two features of $\mathcal{A}_h$ matter for what follows. First, $D_h$ and $C_h$ are functionals of the observed $(p,X)$ alone: $\mathcal{A}_h$ can be estimated without labels and without invoking Assumption~\ref{ass:cal}, which enters only in reading $\mathcal{A}_h$ as the distortion of $\tau$. Second, when $K=1$, \eqref{eq:atten} reduces to the scalar projection slope $\kappa=\mathbb{E}[a_{h}a]/\mathbb{E}[a_{h}^2]$, the familiar dilution factor of a binary pseudo-label, which under a calibrated probability and monotone thresholding is found in $(0,1)$ without being bounded there in general \citep{kurbucz2026kbs}. \citet{chen2019} analyze the same choice for a binary group-mean disparity and decompose the bias of the thresholded estimator, which for that estimand may inflate as well as dilute; the target here is a structural coefficient, and \eqref{eq:atten} collects the whole distortion into the single matrix $\mathcal{A}_h$. What is new with several classes is that the distortion has a direction.

\section{Coarsening and the Validity of Downstream Inference}
\label{sec:coverage}

Theorem~\ref{thm:atten} concerns the point estimate. The more consequential question is what happens to the interval an analyst reports alongside it, because that interval is what enters a decision.

Consider the analyst who computes $\hat\tau_h$, forms the sandwich as though the moment were correctly specified,
\begin{equation}
  \widehat\Omega_h=\widehat D_h^{-1}\widehat\Sigma_h\widehat D_h^{-1},
  \qquad
  \widehat\Sigma_h=n^{-1}\textstyle\sum_i \hat u_{h,i}^{2}\,\hat a_{h,i}\hat a_{h,i}^{\!\top},
  \quad \hat u_h=\hat R-\hat a_h^{\!\top}\hat\tau_h,
  \label{eq:naive-sandwich}
\end{equation}
and reports $v^{\!\top}\hat\tau_h\pm z_{1-\alpha/2}\sqrt{v^{\!\top}\widehat\Omega_h v/n}$ for a direction $v$. Write the directional coarsening bias and variance as $b_v=v^{\!\top}(\mathcal{A}_h-I)\tau$ and $\sigma^2_{h,v}=v^{\!\top}\Omega_h v$.

\begin{proposition}[Coverage under coarsening]
\label{prop:coverage}
Under the conditions of Theorem~\ref{thm:clt} applied to the moment $a_h(R-a_h^{\!\top}\vartheta)$, whose population solution is $\mathcal{A}_h\tau$, along sequences with $b_v=b/\sqrt n$ for fixed $b$,
\begin{equation}
  \mathbb{P}\bigl(v^{\!\top}\tau \in \mathrm{CI}\bigr)\;\longrightarrow\;
  \Phi(z_{1-\alpha/2}-\lambda_v)-\Phi(-z_{1-\alpha/2}-\lambda_v),
  \qquad \lambda_v=b/\sigma_{h,v}.
  \label{eq:coverage}
\end{equation}
For fixed $b_v\neq0$, $\lambda_v=\sqrt n\,b_v/\sigma_{h,v}\to\infty$ and the coverage tends to zero.
\end{proposition}

The proof (in the supplement) is the studentized form of Theorem~\ref{thm:clt} with the probability limit displaced from $\tau$ to $\mathcal{A}_h\tau$. The formula itself is the power function of a $z$-test; what it delivers here is the following.

\begin{remark}[Greater apparent precision can worsen coverage]
\label{rem:precise}
The index $\lambda_v=\sqrt n\,b_v/\sigma_{h,v}$ increases in $1/\sigma_{h,v}$. Coverage therefore degrades faster the more precise the coarsened estimator appears. Coarsened labels are more extreme than the probabilities they replace, so $\Omega_h=D_h^{-1}\Sigma_hD_h^{-1}$ can fall below $\Omega$, as it does in every design below: the reported interval is narrower than the one the uncoarsened regression would have produced, around a displaced center.
\end{remark}

Because $\mathcal{A}_h$ and $\Omega_h$ depend on observables alone, only $b_v$ requires $\tau$, and the uncoarsened estimator supplies it. The coverage the interval will have can therefore be approximated before it is reported:
\begin{equation}
  \hat\lambda_v=\frac{v^{\!\top}(\widehat{\mathcal{A}}_h-I)\hat\tau}{\sqrt{v^{\!\top}\widehat\Omega_h v/n}},
  \qquad
  \widehat{\mathrm{cov}}=\Phi(z_{1-\alpha/2}-\hat\lambda_v)-\Phi(-z_{1-\alpha/2}-\hat\lambda_v).
  \label{eq:cov-hat}
\end{equation}

Three neighboring literatures should be distinguished. \citet{battaglia2024} study regressions on variables generated by machine learning and show that treating them as data yields biased estimates and invalid inference; there the distortion originates in estimation error in the generated variable, and correcting it requires auxiliary information about that error. Here the probability vector is observed exactly and the distortion is a deterministic function of it, which is why \eqref{eq:cov-hat} needs no auxiliary information. Prediction-powered inference \citep{angelopoulos2023} restores validity using gold labels on part of the analysis sample; the diagnostic above is for settings where no such labels exist, and it repairs nothing---it reports what the interval is worth. \citet{li2026} dispenses with assumptions on the upstream classifier altogether, treating the proxy as a link to an auxiliary validation sample and bounding the parameter by optimal transport. That paper also observes that a retained probability vector is more informative than a categorical label, which is the statement \eqref{eq:atten} makes quantitative---again without a validation sample.

\section{Simulation Studies}
\label{sec:sim}

We report here the three experiments most central to the theory, labeled E11, E12 and E6 in a numbering shared with the supplement and the replication code. The data-generating process has $X\in\mathbb{R}^3$ standard normal; a probability vector $p$ drawn from a Dirichlet centred at the softmax mean $\mathrm{softmax}(B_r^{\!\top} W)$ and a latent class $G\mid p\sim\mathrm{Categorical}(p)$ on $\{1,\dots,K+1\}$ ($K=3$), where $W=(X,S)$ augments $X$ with an auxiliary signal $S\in\mathbb{R}^2$ not in the downstream control set; the dispersion parameter $\sigma_u$ controls $\operatorname{Var}(p\mid X)$ and hence the spectrum of $M$; and $Y=X^{\!\top} b_m+g^{\!\top}\tau+\mathcal{N}(0,1)$ with $\tau=(1.0,0.7,-0.4)^{\!\top}$. Nuisances $m,r$ are cross-fitted (5 folds) by degree-2 polynomial regression. Full replication code accompanies the paper. Unless noted, results use $200$ Monte Carlo replicates.

\subsection{The coarsening operator and what it costs (E11)}
Five coarsenings of the same probability vector are compared at $n=8000$: the identity, rounding $p$ to a $0.20$ grid and renormalizing, a top-2 renormalization, a confidence threshold at $0.60$ below which the unit receives the zero vector rather than being dropped, and the argmax one-hot. For each we estimate $\mathcal{A}_h$ from $(p,X)$ alone, run the coarsened estimator, and record the interval it reports along the weakest eigen-direction $v_1$ of $\widehatM$. Table~\ref{tab:e11} shows the pattern the theory predicts. For none of the four non-trivial coarsenings here is the operator a multiple of the identity: its diagonal falls as the coarsening discards more, and its off-diagonal entries are non-zero throughout, so the distortion has a direction. The reported interval meanwhile \emph{narrows} as the coarsening becomes more severe, while its coverage falls from nominal to almost nothing. The argmax interval is $39\%$ shorter than the one the uncoarsened regression would have produced and covers $1\%$ of the time.

\begin{table}[t]
\centering
\spacingset{1}
\caption{Experiment E11: five coarsenings of the same probability vector ($n=8000$, $K=3$,
$400$ replicates). $\mathcal{A}_h$ is estimated from $(p,X)$ alone. SE and coverage
refer to the nominal-$95\%$ Wald interval along the weakest eigen-direction of
$\widehatM$.}
\label{tab:e11}
\begin{threeparttable}
\begin{tabular}{lccrr}
\toprule
Coarsening $h$ & $\operatorname{diag}(\mathcal{A}_h)$ & $\max|\text{off-diag}|$ &
mean SE & coverage\\
\midrule
identity ($h=p$)        & $1.00\ \ 1.00\ \ 1.00$ & $0.000$ & $0.0405$ & $0.952$\\
round to $0.20$ grid    & $0.89\ \ 0.89\ \ 0.82$ & $0.029$ & $0.0377$ & $0.910$\\
top-2 renormalized      & $0.73\ \ 0.75\ \ 0.57$ & $0.059$ & $0.0336$ & $0.675$\\
threshold $0.60$        & $0.47\ \ 0.48\ \ 0.35$ & $0.192$ & $0.0332$ & $0.533$\\
argmax one-hot          & $0.41\ \ 0.42\ \ 0.30$ & $0.096$ & $0.0246$ & $0.010$\\
\bottomrule
\end{tabular}
\end{threeparttable}
\end{table}

\subsection{The coverage formula is calibrated (E12)}
Proposition~\ref{prop:coverage} predicts coverage from the noncentrality index $\lambda_v$. Grouping replicates by $|\lambda_v|$ and comparing realized with predicted coverage within groups, the formula tracks the realized value to within $0.02$ across the whole range, from nominal coverage down to $0.003$. Substituting $\hat\tau$ for $\tau$ in $\hat\lambda_v$ widens these gaps, most where the distortion is most severe, so \eqref{eq:cov-hat} is best read as an accurate ordering and an approximate level. A sweep over the dispersion of $p$ and over the overlap between the classifier's features and the controls leaves the ordering unchanged: in every configuration examined the argmax interval covers between $1\%$ and $6\%$. The grouped comparison and the sweep are tabulated in the supplement.

\subsection{Inference for the uncoarsened estimator (E6)}
Componentwise and joint Wald coverage for $\hat\tau$ itself lie in $[0.938,0.960]$ and $[0.946,0.958]$ at nominal $95\%$ over $1000$ replicates as $n$ grows from $500$ to $5000$, consistent with Theorem~\ref{thm:clt}; the table and the normal QQ-plot are in the supplement. The contrast with Table~\ref{tab:e11} is the point of the paper: the same design, the same nominal level, and coverage of $0.95$ or of $0.01$ along the same contrast, according to whether the probability vector was coarsened first.

Four further experiments, reported in the supplement, complete the picture. Argmax coarsening acts anisotropically, and in that design the retention coefficient increases with the eigenvalue, one instance of the directional heterogeneity of Corollary~\ref{cor:aniso}. Setting $K=1$ recovers the scalar $(V^{*})^{-2}$ sandwich law. The efficient instrument of Theorem~\ref{thm:eff} cuts the variance threefold, while the naive weighting that omits its recentering reports standard errors that are too small. Finally, inference remains valid at the $\sqrt{n\lambda}$ rate along a collapsing direction.

\section{A Disparity Audit with a Proxied Protected Attribute}
\label{sec:adult}

The motivating application of Section~\ref{sec:intro}---auditing outcome disparities across groups when the protected attribute is unobserved and must be proxied \citep{kallus2022,chen2019}---is naturally multi-class. We use the standard training file of UCI Adult \citep{kohavi1996} ($n=30{,}162$ after listwise deletion). The latent group is race in $K+1=4$ classes (reference \emph{White}; classes \emph{Black}, \emph{Asian--Pacific}, \emph{Other}); the controls $X$ are age, education, and sex; the proxy is a multinomial classifier on a richer feature set $W\supsetneq X$ (occupation, workclass, relationship, native region, capital variables), cross-fitted over $5$ folds to an out-of-fold probability vector $p$. A controlled validation on this real score spectrum, in which the estimator recovers a known $\tau$ along every eigen-direction, is reported in the supplement; here we run the genuine audit.

The outcome $Y$ is the high-income indicator, the attribute is treated as unobserved, and the estimator uses only the proxy $p$. Because race is in fact recorded here, we also compute the full-information benchmark from the true labels; the hard estimator uses the argmax class. There is no guarantee the assumptions hold---an off-the-shelf race proxy need not be conditionally calibrated, and its features may carry direct income channels.

The estimated spectrum of $\widehatM$ is anisotropic, with eigenvalues $(0.0006,0.0069,0.018)$ and condition number $31$: one disparity direction is reasonably identified and two are weak. Table~\ref{tab:adult} reports, in the eigenbasis of $\widehatM$, the retention coefficient under argmax coarsening, the estimate from the proxy alone, and the full-information benchmark. The retention coefficients range from $s_1=0.46$ to $s_3=0.90$: on real data, coarsening removes more than half the signal in one direction and almost none in another. This is Corollary~\ref{cor:aniso} outside a simulation, and it requires no calibration assumption, since $\widehat{\mathcal{A}}_h$ is computed from $(p,X)$ alone.

Along the best-identified direction the proxy-based estimate matches the benchmark ($0.034\pm0.017$ against $0.028$); along the two weak directions it departs sharply, as eigenvalues an order of magnitude smaller amplify whatever violation is present. The supplement reports clean calibration batteries but an emphatic direct-channel check ($\Delta R^2=0.135$), which points to the exclusion leg rather than calibration. We report this as a cautionary instance, not a substantive claim about disparities.

\begin{table}[t]
\centering
\spacingset{1}
\caption{UCI Adult disparity audit, in the eigenbasis of $\widehatM$
(weakest to strongest identified direction). $s_j$: retention coefficient under
hard-labeling; soft: proxy-based estimate of $v_j^{\!\top}\tau$ (sandwich SE);
benchmark: same projection from the observed labels.}
\label{tab:adult}
\begin{threeparttable}
\begin{tabular}{lrrrr}
\toprule
Direction & $\lambda_j(\widehatM)$ & $s_j$ &
soft $v_j^{\!\top}\hat\tau$ (SE) & benchmark $v_j^{\!\top}\hat\tau$\\
\midrule
$v_1$ (weakest)   & 0.0006 & 0.46 & $-0.584\ (0.080)$ & $-0.079$\\
$v_2$             & 0.0069 & 0.72 & $-0.472\ (0.025)$ & $-0.065$\\
$v_3$ (strongest) & 0.0177 & 0.90 & $\phantom{-}0.034\ (0.017)$ & $\phantom{-}0.028$\\
\bottomrule
\end{tabular}
\end{threeparttable}
\end{table}


\section{Diagnosing Assumption Failure by Recalibration}
\label{sec:recal}

The framework rests on two substantive legs: conditional calibration (Assumption~\ref{ass:cal}) and the exclusion component of Assumption~\ref{ass:mean}---the requirement that the score carry no direct outcome channel given $(G,X)$. The Adult audit above, and a fully real land-cover audit on UCI Forest CoverType (a distance-to-water outcome with a random-forest cover-type classifier; details in the supplement), both fail somewhere; a practitioner needs to know which leg broke, because only one of them can be repaired on a labeled subsample.

A labeled subsample makes the question answerable by experiment. Split the labeled data in half; on the first half, fit the recalibration map $\tilde p=\widehat{\mathbb{E}}[g\mid p,X]$ (here, a multinomial logistic regression on the log-odds of $p$, the controls, and their interactions); on the held-out half, re-run the audit with $\tilde p$ in place of $p$. Recalibration targets Assumption~\ref{ass:cal}: were the population map known exactly, $\mathbb{E}[g\mid\tilde p,X]=\tilde p$ would hold by the tower property, and the out-of-sample battery measures how nearly the estimated map achieves it. It does not target the exclusion leg: if $p$ carries a direct outcome channel, so may any function of $(p,X)$. And if Assumption~\ref{ass:mean} holds for $(p,X)$ it holds for $(\tilde p,X)$, since $\sigma(\tilde p,X)\subseteq\sigma(p,X)$. The movement of the audit under recalibration is therefore informative about which failure dominates, read together with the calibration battery and the direct-channel check.

The mechanism is one line from the proof of Theorem~\ref{thm:ident}. If the structural mean carries a direct channel, $\mathbb{E}[Y\mid G,p,X]=\mu(X)+\tau^{\!\top} g+d(p,X)$ for a direct channel $d$, then $\mathbb{E}[aR]=M\tau+\mathbb{E}[a\,d]$ and $\operatorname{plim}\hat\tau=\tau+M^{-1}\mathbb{E}[a\,d]$: the violation enters divided by the eigenvalues of $M$. In our audits recalibration compresses the probability vector toward its conditional expectation, shrinking the spectrum while leaving the direct channel intact, so recalibration can amplify the bias from a live exclusion violation.

\begin{proposition}[Sharp sensitivity to miscalibration]
\label{prop:cal}
Let Assumptions~\ref{ass:mean}, \ref{ass:rank} and~\ref{ass:mom} hold, and let $\mathbb{E}[g\mid p,X]=p+\eta(p,X)$ replace Assumption~\ref{ass:cal}, with $\mathcal{H}_\delta$ collecting the errors with $\|\eta\|_\infty\le\delta$ almost surely for which $p+\eta$ remains a valid class-probability vector. Then: (a) $\operatorname{plim}\hat\tau=\tau+M^{-1}N\tau$ with $N=\mathbb{E}[a\,\eta^{\!\top}]$; the bias is linear in $\tau$, so if $\tau=0$ miscalibration alone cannot produce a nonzero coefficient vector, although a single null component can still be distorted by the mixing. (b) For every unit vector $v$,
\begin{equation}
  \sup_{\eta\in\mathcal{H}_\delta}
  \bigl|v^{\!\top}(\operatorname{plim}\hat\tau-\tau)\bigr|
  \;\le\;\delta\,\|\tau\|_1\,\mathbb{E}\bigl|v^{\!\top}M^{-1}a\bigr|,
  \label{eq:cal-bound}
\end{equation}
with equality whenever every class probability, including the reference, is at least $K\delta$ almost surely. Along the eigen-direction $v_j$ the bound equals $\delta\|\tau\|_1\mathbb{E}|v_j^{\!\top} a|/\lambda_j\le\delta\|\tau\|_1\lambda_j^{-1/2}$.
\end{proposition}

Proposition~\ref{prop:cal} (proved in the supplement) is the quantitative form of the warning after Assumption~\ref{ass:cal}: the bound grows as the eigenvalue shrinks and is no larger than $\delta\|\tau\|_1\lambda_j^{-1/2}$, so the directions least identified, and least precisely estimated, are also the most fragile to calibration drift. The proposition is the specialization to this model of the approximate-moment-condition framework of \citet{armstrong2021}, in which a moment restriction is allowed to fail within a budget and the resulting identified set and confidence intervals are derived in general. What the model adds is a closed form---the bound is attained when every class probability is at least $K\delta$---and an eigenvalue-indexed reading of which contrasts a given budget puts beyond reach. For inference on the resulting set rather than the point, that general machinery applies directly. It also informs the protocol below: recalibration shrinks $\delta$ but compresses the spectrum, so its net effect is direction-specific.

Table~S5 in the supplement reports the experiment on both audits, together with an independent check of the exclusion leg (the partial $R^2$ of the classifier features $W$ for the outcome given the true class and $X$). The two audits fail with different dominant failure patterns. On Adult the rich-basis battery finds essentially no conditional miscalibration ($R^2\le0.006$), recalibration accordingly changes nothing, and the direct-channel check is emphatic ($\Delta R^2=0.135$): the classifier's features predict earnings directly, so the failure of the weak directions is an exclusion violation that no calibration step can touch. On CoverType the picture reverses. The rich basis reveals severe conditional miscalibration ($R^2$ up to $0.32$) and recalibration reduces it out of sample to $R^2\le0.003$, yet the audit deteriorates. The recalibrated spectrum compresses (largest eigenvalue $0.122\to0.029$) and the direct channel carried by the soil and wilderness indicators ($\Delta R^2=0.081$) is amplified exactly as the displayed bias expression predicts.


The protocol that emerges is sequential: run the rich-basis $(p,X)$ calibration battery on the labeled subsample, since an $X$-only check can miss violations that live in the $p$-direction; then recalibrate and re-run out of sample. If the gap to an available benchmark closes, the evidence points to calibration as the dominant failure, which is repairable. If the audit barely moves despite a clean battery, or moves sharply after recalibration, the exclusion leg is live and the soft audit remains unreliable under these diagnostics. The experiment requires no labeled data beyond the subsample the battery already uses.

\section{A Voter-Turnout Disparity Audit}
\label{sec:bisg}

The audits so far have been honest about failure: on UCI Adult and CoverType the exclusion leg is live. This is a property of the data rather than of the method: in a generic observational regression the classifier's features may carry outcome-relevant information beyond the latent class. The framework is designed for a different structure, a measurement model in which the classifier reads a proxy that is a pure indicator of the latent class, with no direct channel to the outcome. We now exhibit a fully real instance in which the diagnostics are favorable on every operating condition, though conditional calibration is not exact.

We audit racial disparities in voter turnout when the protected attribute is proxied rather than observed, the setting of the surname- and geography-based race-proxy literature \citep{elliott2009,imai2016}. The data are the public North Carolina voter file and its voter-history file for Mecklenburg County ($n=464{,}700$ active registrants after matching), together with the 2010 Census surname list \citep{comenetz2016}. The latent group $G$ is self-reported race/ethnicity in $K+1=4$ classes (reference \emph{White}; \emph{Black}, \emph{Hispanic}, \emph{Asian}), recorded on the voter file and used only to form the benchmark. The probability vector $p$ is the surname list's national $\mathbb{P}(\text{race}\mid\text{surname})$, an off-the-shelf classifier we take as given; the controls $X$ are the electoral precinct; the outcome $Y$ is the number of November general elections in which the registrant voted, 2016--2024. Unlike Bayesian Improved Surname Geocoding, which folds geography into the proxy itself, we keep the proxy surname-only and let precinct enter separately as a control. The framework needs $X$ strictly coarser than the classifier's features, and a proxy already conditioned on geography would leave little residual variation once precinct is partialed out.

A surname is a plausible \emph{pure race indicator}: it carries genuine, moderate information about race, while any direct channel to turnout once race and place are fixed is hard to motivate substantively. The data are consistent with this. The direct-channel check of Section~\ref{sec:recal} is essentially zero: the partial $R^2$ of the surname score for turnout given true race and precinct is $0.001$. The surname classifier is genuinely uncertain but skilful (argmax accuracy $0.63$ against a White base rate of $0.55$; many surnames are racially ambiguous), so the score is neither degenerate nor deterministic. And the residual-score covariance $\widehatM$ is well conditioned (eigenvalues $0.015,0.024,0.061$; condition number $4.1$): no direction appears close to collapse.

Table~\ref{tab:bisg} reports the estimated turnout gaps. The full-information benchmark (from self-reported race) shows Hispanic and Asian registrants turning out roughly half a standard deviation below White registrants net of precinct, and Black registrants modestly below. Hard-labeling attenuates these gaps toward zero---consistent with the direction-specific distortion of Theorem~\ref{thm:atten}---and does most damage where the true gap is largest, shrinking the Hispanic and Asian gaps by roughly a fifth ($-0.454$ and $-0.428$ against benchmark values $-0.551$ and $-0.525$). The soft estimator, built from the surname proxy alone, recovers them almost exactly ($-0.520$ and $-0.524$). Overall the soft estimator is twice as close to the benchmark as hard-labeling (RMSE $0.042$ versus $0.082$), a gap stable across resamplings.

\begin{table}[t]
\centering
\spacingset{1}
\caption{Census-surname voter-turnout audit, North Carolina (Mecklenburg County,
$n=464{,}700$), in the original class basis. Benchmark uses self-reported
race; soft and hard use only the surname proxy. Entries are turnout gaps
relative to White, net of precinct (outcome standardized).}
\label{tab:bisg}
\begin{threeparttable}
\begin{tabular}{lrrrr}
\toprule
Group & benchmark (true race) & soft (surname) & hard (argmax) & hard/benchmark\\
\midrule
Black    & $-0.076$ & $-0.011$ & $-0.040$ & $0.53$\\
Hispanic & $-0.551$ & $-0.520$ & $-0.454$ & $0.82$\\
Asian    & $-0.525$ & $-0.524$ & $-0.428$ & $0.82$\\
\bottomrule
\end{tabular}
\begin{tablenotes}\small
\item \textit{Notes}: RMSE versus the benchmark: soft $0.042$, hard $0.082$.
Direct-channel check (exclusion) $\Delta R^2=0.001$; surname-argmax accuracy
$0.63$ (White base rate $0.55$); $\widehatM$ condition number $4.1$. The last column is the ratio of the hard-label estimate to the benchmark, in the original class basis; it is not the eigen-direction retention coefficient $s_j$.
\end{tablenotes}
\end{threeparttable}
\end{table}

Two caveats keep the claim precise. At $n\approx465{,}000$ the uncoarsened estimator does not match the benchmark exactly: its small deviations are statistically resolvable, and on the already-small Black gap the residual bias makes it no better than the coarsened one. This reflects that a national surname list is not exactly conditionally calibrated to a single county; the point is that it is right about the magnitudes of the large gaps that coarsening systematically shrinks. The gain is also not an artifact of recalibration: the raw Census probabilities already win, and recalibrating them locally on a labeled half leaves the conclusion unchanged.

\section{Discussion}
\label{sec:disc}

Two objects organize the analysis. The covariance matrix $M$ of the partialed-out probability vector grades which contrasts the regression \eqref{eq:plr} can recover and how precisely; the coarsening operator $\mathcal{A}_h$ grades what is lost when that vector is replaced by a coarser version of itself, and, through the noncentrality index, what the reported interval is then worth. The binary case is the one-dimensional shadow of both: there the operator is a scalar dilution factor and identification is all or nothing, while with several classes each becomes direction-specific.

\subsection*{Scope and operating conditions}
The three real-data audits delimit the operating range of the framework, which is informative when all of the following hold. (i)~Membership in several mutually exclusive latent classes is unobserved, and the probability vector is conditionally calibrated. This holds by construction for a correctly specified posterior built on information strictly richer than the controls, and it is testable, and improvable by recalibration, when a labeled subsample is available (Section~\ref{sec:recal}). The CoverType audit shows what happens when an off-the-shelf classifier is deployed without that step. (ii)~Whatever controls $X$ the outcome equation requires are strictly coarser than the classifier's features, so that residual variation in the score survives partialling out: without it the no-controls estimator in the land-cover audit (supplement) sign-flips, and with $X$ as rich as $W$ the residual $a$ degenerates and the spectrum collapses. (iii)~The classifier is genuinely uncertain. When it is nearly deterministic the probability vector concentrates on the vertices of the simplex, coarsened and uncoarsened labels coincide, $\mathcal{A}_h\to I$, and the framework is valid but adds nothing: the gain is largest exactly for hard-but-learnable tasks, where coarsening is also most destructive. (iv)~There is no trusted forward model of the measurement process; where a generative likelihood for the probability vector is available and believed, full-likelihood methods may be more efficient.

Settings that satisfy all four conditions are the natural applications: multi-group fairness audits with validation labels, hard medical subtyping against registry gold standards, verbal-autopsy cause-of-death attribution, and multi-subtype prevalence estimation with administratively missing indicators. The surname-proxy turnout audit of Section~\ref{sec:bisg} comes closest to satisfying all four. Land-cover and similar remote-sensing problems qualify only conditionally. The contrast is structural: the favorable case uses a pure-indicator proxy (a surname carries race but does not cause turnout), whereas the conditional cases use covariates that also drive the outcome. Where a labeled subsample exists, the recalibration experiment of Section~\ref{sec:recal} upgrades this checklist from a judgment call to a measurement: it tests the calibration leg and indicates which of the two failures dominates.

Three directions follow. The efficient estimator attains the bound but needs two nuisances beyond $(m,r)$, the conditional variance and the precision-weighted mean; the supplement shows the gain survives a well-specified variance model and is lost under a poor one, and behavior under fully nonparametric estimation of both is open. Inference on the set induced by a calibration budget can be carried out with the machinery of \citet{armstrong2021}, which is built for moment conditions that fail by a bounded amount. The identified set along a collapsed direction is a different problem: by Theorem~\ref{thm:spectral}(d) it does not shrink with the sample size. Third, the coarsening operator is the core of an applied diagnostic: a direction-by-direction rule for choosing among the label-based benchmark and the uncoarsened and coarsened estimators. For a binary latent class \citet{kurbucz2026kbs}\ develops such a rule, indexing the attenuation factor by the confidence threshold; carrying it over to several classes is left to future work.