EconBase
← Back to paper

Identification of Latent Group Effects under Conditional Calibration

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

41,098 characters

Identification of Latent Group Effects under Conditional Calibration



\begin{frontmatter}

\title{Identification of Latent Group Effects under Conditional
  Calibration}

\author{Marcell T.\ Kurbucz\corref{cor1}}
\ead{[email removed]}
\cortext[cor1]{Corresponding author.}
\affiliation{organization={Institute for Global Prosperity, The Bartlett,
               University College London},
             addressline={9--11 Endsleigh Gardens},
             city={London},
             postcode={WC1H 0EH},
             country={United Kingdom}}

\begin{abstract}
\noindent
\footnotesize
We study identification of a structural group effect when the group indicator
$G\in\{0,1\}$ is unobserved but the analyst observes a calibrated probability
score $p\in[0,1]$ satisfying $\mathbb{E}[G\mid p,X]=p$. Under a constant-coefficient
structural mean model, the latent-group coefficient $\tau$ is point-identified
from the joint law of observables $(Y,X,p)$ by a simple ratio of weighted
moments: the covariance of the signed score $2p-1$ with the
covariate-partialled outcome, divided by twice the residual variance of the
score after conditioning on covariates. Identification fails if and only if
the score is a deterministic function of $X$; we establish this by constructing
an explicit continuum of observationally equivalent models indexed by arbitrary
values of $\tau$. The identified coefficient differs from the marginal latent
mean gap by a compositional term that is unidentified without further
assumptions; we give a necessary and sufficient condition for the two to
coincide. The oracle estimator is $\sqrt{n}$-consistent and asymptotically
normal with a closed-form sandwich variance. Under calibration error bounded
uniformly by $\delta$, the bias is bounded by
$|\tau|\,\mathbb{E}[|2p-1|]\,\delta\,(2V^*)^{-1}$, a bound that is sharp over all
calibration error functions of that magnitude. Hard-threshold classification
at $p=\tfrac{1}{2}$ attenuates the estimated gap by a factor strictly less than
one. Monte Carlo experiments confirm the asymptotic theory, trace the
divergence of RMSE as $V^*\to 0$, illustrate the attenuation bias of
hard-threshold classification, and verify identification of the
variance-weighted estimand under heterogeneous effects.
\end{abstract}

\begin{keyword}
\footnotesize
latent groups \sep identification \sep conditional calibration \sep
moment equations \sep group effects \\[4pt]
\textit{JEL codes:} C14 \sep C21 \sep C38 \sep D63
\end{keyword}

\end{frontmatter}

\section{Introduction}
\label{sec:intro}
\noindent
A pervasive challenge in empirical work is the measurement of outcome
differences between groups when group membership is not directly observed.
Poverty status, immigration status, informal employment, fuel insecurity, and
latent health conditions are leading examples. In such settings the analyst
typically has access to a probability score $p_i\in[0,1]$ encoding belief
that unit $i$ belongs to the group of interest, but never observes the binary
indicator $G_i\in\{0,1\}$ itself.

The central question we address is: \emph{under what conditions, and by what
formula, can a structural group effect be identified from the joint law of
observables $(Y,X,p)$ when $G$ is never observed?}  We give a precise
answer organised around three claims. First, a structural coefficient $\tau$
is point-identified under mild conditions. Second, identification fails in a
characterisable and sharp way when exactly one of those conditions is violated.
Third, the identified object is distinct from the marginal group mean gap in a
way that can be made fully explicit.

The paper makes four contributions. The first is an identification result.
Under a constant-coefficient structural mean model and the conditional
calibration condition $\mathbb{E}[G\mid p,X]=p$, we prove that $\tau$ is identified
by a weighted moment equation whose denominator
$V^{*}=\mathbb{E}[(p-r(X))^{2}]$ is the residual variance of the score after partialling
on $X$. The formula is in closed form and admits a transparent interpretation:
it is formally analogous to an instrumental-variables estimand in which the
score residual $a=p-r(X)$ plays the role of an instrument for the latent
deviation $G-r(X)$; the calibration condition supplies the first-stage
relevance and the mean-independence condition in the structural model supplies
the exclusion restriction.

The second contribution is an exact characterisation of identification failure.
We prove that $\tau$ is not identified if and only if $V^{*}=0$, i.e.\ the score
is a deterministic function of $X$, and we make this concrete by constructing
an explicit continuum of observationally equivalent models indexed by arbitrary
values $\tau'\in\mathbb{R}$.

The third contribution is a clean separation between the identified structural
coefficient and the marginal latent mean gap
$\Delta_{\mathrm{marg}}=\mathbb{E}[Y\mid G=1]-\mathbb{E}[Y\mid G=0]$. We decompose
$\Delta_{\mathrm{marg}}=\tau+C$ where the compositional term $C$ is not
identified from $(Y,X,p)$ alone, and we give a necessary and sufficient
condition for $C=0$.

The fourth contribution is oracle inference and robustness. We establish
$\sqrt{n}$-asymptotic normality of the oracle estimator with an explicit
sandwich variance, compute the exact probability limit under calibration
failure, and derive a sensitivity bound that is sharp over the class of all
calibration error functions bounded uniformly by $\delta$.

Our paper sits at the intersection of four strands of the literature.
Within the \emph{misclassification} literature, \citet{lewbel2007} showed that
average treatment effects are attenuated under misclassification of a binary
regressor and proposed corrections; \citet{mahajan2006} obtained identification
using an instrumental variable; \citet{kasahara2022} extended this to the
endogenous case. Our setting is complementary: instead of observing a noisy
binary label, the analyst observes a calibrated \emph{probability} for $G$,
which changes both the identification argument and the identified object.

The \emph{proxy variable and measurement error} literature
\citep{hu2008,schennach2016} establishes nonparametric identification of full
latent-variable distributions via rank conditions on integral operators. Our
setting is more restrictive---we target only the scalar $\tau$---but our
assumptions are correspondingly weaker and the identification formula is
closed-form. The structure of our moment equation parallels the partially
linear model \citep{robinson1988} and semiparametric IV \citep{newey1990};
the key departure is that the ``instrument'' is not an observable variable
but a residual derived from a calibrated probability. Finally, the
\emph{algorithmic fairness} literature \citep{kallus2022,chen2019} has studied
disparity estimation with unobserved protected attributes under
calibration-type assumptions; our contribution to that context is a formal
identification-theoretic treatment with a closed-form formula and exact failure
and sensitivity characterisations.

The remainder of the paper is organised as follows.
Section~\ref{sec:model} sets up the model.
Section~\ref{sec:ident} proves identification and characterises failure.
Section~\ref{sec:structural} distinguishes the structural coefficient from
the marginal gap.
Section~\ref{sec:inference} covers oracle inference and the plug-in
estimator, including Neyman orthogonality.
Section~\ref{sec:robustness} develops robustness to calibration failure.
Section~\ref{sec:mc} presents Monte Carlo evidence.
Section~\ref{sec:discussion} discusses extensions and open directions.
Section~\ref{sec:conclusion} concludes.
The Appendix contains all proofs.

\section{Model and Assumptions}
\label{sec:model}
\noindent
Let $(\Omega,\mathcal{F},\mathbb{P})$ be a probability space. We observe i.i.d.\
draws $(Y_i,X_i,p_i)\in\mathbb{R}\times\mathcal{X}\times[0,1]$ for $i=1,\ldots,n$,
from the joint distribution $P_{Y,X,p}$. There exists on the same probability
space an unobserved binary variable $G_i\in\{0,1\}$; $G_i=1$ denotes
membership in the latent group of interest. The measurable covariate space
$\mathcal{X}$ is arbitrary.

We write $m(x):=\mathbb{E}[Y\mid X=x]$, $r(x):=\mathbb{E}[p\mid X=x]$, and
$\pi(x):=\mathbb{E}[G\mid X=x]$ for the conditional mean functions. The derived
quantities
\begin{equation}
  z:=2p-1\in[-1,1],\quad R:=Y-m(X),\quad a:=p-r(X)
  \label{eq:derived}
\end{equation}
satisfy $\mathbb{E}[R\mid X]=\mathbb{E}[a\mid X]=0$. The \emph{residual score variance}
\begin{equation}
  V^{*}:=\mathbb{E}\!\left[(p-r(X))^{2}\right]=\mathbb{E}[\operatorname{Var}(p\mid X)]\;\geq 0
  \label{eq:Vs}
\end{equation}
measures the variation in $p$ not explained by $X$; it is the key quantity
governing identification.

\begin{assumption}[Structural conditional mean]
\label{ass:mean}
There exist a measurable function $\mu:\mathcal{X}\to\mathbb{R}$ and a scalar
$\tau\in\mathbb{R}$ such that $\mathbb{E}[Y\mid G,p,X]=\mu(X)+\tau G$ a.s.
\end{assumption}

Assumption~\ref{ass:mean} has two components. The effect of latent membership
on the conditional mean of $Y$ is constant in $X$. Additionally, the score
$p$ is mean-independent of $Y$ once $(G,X)$ are known: conditional on true
membership, the analyst's probability score conveys no further information
about the expected outcome.

\begin{assumption}[Conditional calibration]
\label{ass:cal}
$\mathbb{E}[G\mid p,X]=p$ a.s.
\end{assumption}

Assumption~\ref{ass:cal} is the sole formal link between the latent indicator
$G$ and the observed score $p$. It is a calibration condition: $p$ need not
equal the propensity score $\mathbb{P}(G=1\mid X)$ but must be an unbiased predictor
of $G$ given all observed information $(p,X)$. Scores arising from area-level
prevalence rates, classifier outputs, or model-based predictions all satisfy
this condition when they are well-calibrated.

\begin{assumption}[Non-degenerate residual variation]
\label{ass:var}
$V^{*}>0$.
\end{assumption}

\begin{assumption}[Moment conditions]
\label{ass:moments}
$\mathbb{E}[Y^{4}]<\infty$ and $\mathbb{E}[p^{4}]<\infty$.
\end{assumption}

Assumption~\ref{ass:moments} implies square-integrability of all relevant
quantities and is used in Sections~\ref{sec:ident}--\ref{sec:inference}.
Asymptotic normality in Section~\ref{sec:inference} uses the full fourth-moment
condition; identification requires only second moments.

The following two lemmas are the structural backbone of the paper.

\begin{lemma}[Structural decomposition]
\label{lem:basic}
Under Assumptions~\ref{ass:mean}--\ref{ass:moments},
\begin{equation}
  \pi(X)=r(X)\text{ a.s.}, \qquad m(X)=\mu(X)+\tau r(X)\text{ a.s.}
  \label{eq:struct-decomp}
\end{equation}
\end{lemma}

\begin{lemma}[Residual decomposition]
\label{lem:residual}
Under Assumptions~\ref{ass:mean}--\ref{ass:moments}, the outcome residual
satisfies
\begin{equation}
  R=\tau(G-r(X))+\varepsilon,\qquad \mathbb{E}[\varepsilon\mid G,p,X]=0.
  \label{eq:resid-decomp}
\end{equation}
\end{lemma}

Proofs are in Appendix~\ref{app:proofs}. The content of
Lemma~\ref{lem:residual} is structural: the entire predictable part of the
outcome residual $R$ is driven by the deviation of true group membership from
its score-implied expectation, $G-r(X)$.

\section{Identification}
\label{sec:ident}
\noindent
\begin{theorem}[Population moment identity]
\label{thm:moment}
Under Assumptions~\ref{ass:mean}--\ref{ass:moments},
\begin{equation}
  \mathbb{E}\!\left[(2p-1)(Y-m(X))\right]=2\tauV^{*}.
  \label{eq:moment}
\end{equation}
\end{theorem}

\begin{corollary}[Point identification]
\label{cor:ident}
Under Assumptions~\ref{ass:mean}--\ref{ass:var},
\begin{equation}
  \tau=\frac{\mathbb{E}[(2p-1)(Y-m(X))]}{2\,\mathbb{E}[(p-r(X))^{2}]}.
  \label{eq:tau-id}
\end{equation}
The structural coefficient $\tau$ is point-identified from the joint law of
$(Y,X,p)$.
\end{corollary}

The proof of Theorem~\ref{thm:moment} is in Appendix~\ref{app:thm-moment}.
The identification formula~\eqref{eq:tau-id} has a transparent algebraic
structure. The numerator $\mathbb{E}[zR]$ is the covariance between the signed score
$z=2p-1$ and the outcome residual $R$, after partialling both on $X$. The
denominator $2V^{*}=2\mathbb{E}[a^{2}]$ is twice the residual variance of the score.
The ratio is therefore the slope of the regression of $R$ on $z$ in the
covariate-partialled data, which by Lemma~\ref{lem:cov-gp} in the Appendix
equals the slope of the regression of $R$ on the latent deviation $G-r(X)$.
This is formally analogous to an IV estimand in which $a=p-r(X)$ acts as an
instrument for $G-r(X)$: the calibration condition (Assumption~\ref{ass:cal})
supplies the first-stage relevance, and the mean-independence condition in
Assumption~\ref{ass:mean} supplies the exclusion restriction.

We now show that Assumption~\ref{ass:var} is not merely a regularity
condition but the exact boundary of identification.

\begin{proposition}[Identification failure]
\label{prop:failure}
\begin{enumerate}[label=\textup{(\alph*)}]
  \item If $V^{*}=0$, then both sides of~\eqref{eq:moment} equal zero for
    every $\tau\in\mathbb{R}$.
  \item If $V^{*}=0$, then for every $\tau'\in\mathbb{R}$ there exists a model
    satisfying Assumptions~\ref{ass:mean} and~\ref{ass:cal} with
    latent-group coefficient $\tau'$ that implies the same conditional
    mean $m(X)=\mathbb{E}[Y\mid X]$ as the original model and hence the same
    moment equation~\eqref{eq:moment}.
  \item $V^{*}=0$ if and only if $p=r(X)$ almost surely.
\end{enumerate}
\end{proposition}

Part (b) is the substantive non-identification claim. The explicit
construction is in Appendix~\ref{app:prop-failure}. It defines a new model
$\mathcal{M}'$ that retains the observable distribution of $(Y,X,p)$,
sets $G':=\mathbf{1}\{U\leq p\}$ for an independent $U\sim\mathrm{Uniform}(0,1)$,
and postulates the structural equation with coefficient $\tau'$ and
$\mu'(X):=m(X)-\tau'r(X)$. The postulate is consistent because the implied
observable regression of $Y$ on $X$ equals $m(X)$ for every $\tau'\in\mathbb{R}$,
so $\mathcal{M}'$ satisfies both assumptions while being observationally
indistinguishable from the original model.

\section{The Structural Coefficient and the Marginal Gap}
\label{sec:structural}
\noindent
A natural question is whether $\tau$ equals the marginal latent mean gap
$\Delta_{\mathrm{marg}}:=\mathbb{E}[Y\mid G=1]-\mathbb{E}[Y\mid G=0]$. Under
Assumption~\ref{ass:mean}, a direct calculation gives
\begin{equation}
  \Delta_{\mathrm{marg}}=\tau+C,\qquad
  C:=\mathbb{E}[\mu(X)\mid G=1]-\mathbb{E}[\mu(X)\mid G=0].
  \label{eq:marg-decomp}
\end{equation}
The term $C$ captures differences in covariate composition across latent
groups. It depends on the latent conditional distributions $\mathbb{P}_{X\mid G=g}$,
which are not identified from $(Y,X,p)$ under our assumptions: many latent
joint distributions of $(G,X)$ generate the same observable distribution.

\begin{proposition}[Structural coefficient versus marginal gap]
\label{prop:sv-marginal}
Under Assumptions~\ref{ass:mean}--\ref{ass:var}, the following are
equivalent: (i)~$\Delta_{\mathrm{marg}}=\tau$; (ii)~$C=0$;
(iii)~the latent groups are covariate-balanced,
$\mathbb{E}[\mu(X)\mid G=1]=\mathbb{E}[\mu(X)\mid G=0]$.
\end{proposition}

The proof is immediate from~\eqref{eq:marg-decomp}. The practical import is
that $\tau$ identifies the \emph{within-covariate-cell} group effect, while
$\Delta_{\mathrm{marg}}$ conflates this with the compositional term $C$.
Recovering $\Delta_{\mathrm{marg}}$ requires identifying $C$, which in turn
requires knowledge of the latent-group covariate distributions---not available
under our assumptions without further restrictions.

\section{Estimation and Inference}
\label{sec:inference}

\subsection{Oracle estimator}
\label{sec:oracle}
\noindent
Suppose for this subsection that $m$ and $r$ are known. Define the
\emph{oracle estimator}
\begin{equation}
  \hat{\tau}_{\mathrm{or}}:=\frac{\frac{1}{n}\sum_{i}(2p_i-1)(Y_i-m(X_i))}
               {2\,\frac{1}{n}\sum_{i}(p_i-r(X_i))^{2}},
  \label{eq:oracle}
\end{equation}
and the score evaluated at the true parameter,
\begin{equation}
  \psi_i:=(2p_i-1)(Y_i-m(X_i))-2\tau(p_i-r(X_i))^{2}.
  \label{eq:score}
\end{equation}
By Theorem~\ref{thm:moment}, $\mathbb{E}[\psi_i]=0$.

\begin{theorem}[Oracle CLT]
\label{thm:clt}
Under Assumptions~\ref{ass:mean}--\ref{ass:moments} and i.i.d.\ sampling,
\begin{equation}
  \sqrt{n}\,(\hat{\tau}_{\mathrm{or}}-\tau)\;\xrightarrow{\;d\;}\;\mathcal{N}(0,\,\sigma^{2}_{\mathrm{or}}),
  \qquad
  \sigma^{2}_{\mathrm{or}}
  =\frac{\mathbb{E}[\psi_i^{2}]}{(2V^{*})^{2}}.
  \label{eq:clt}
\end{equation}
\end{theorem}

The variance $\sigma^{2}_{\mathrm{or}}=J^{-2}\mathbb{E}[\psi_i^{2}]$ has the
standard sandwich form with Jacobian $J=2V^{*}$. The identification
condition $V^{*}>0$ is precisely the condition that $J\neq 0$, i.e., that the
moment equation is locally informative about $\tau$ in a neighbourhood of the
truth.

\begin{proposition}[Consistent variance estimator and Wald interval]
\label{prop:wald}
Let $\hat\psi_i:=(2p_i-1)(Y_i-m(X_i))-2\hat{\tau}_{\mathrm{or}}(p_i-r(X_i))^{2}$. The
estimator
\begin{equation}
  \hat\sigma^{2}_{\mathrm{or}}
  :=\frac{\frac{1}{n}\sum_{i}\hat\psi_i^{2}}
         {\Bigl(2\,\frac{1}{n}\sum_{i}(p_i-r(X_i))^{2}\Bigr)^{2}}
  \label{eq:varhat}
\end{equation}
satisfies $\hat\sigma^{2}_{\mathrm{or}}\xrightarrow{\;p\;}\sigma^{2}_{\mathrm{or}}$, and
$\hat{\tau}_{\mathrm{or}}\pm z_{1-\alpha/2}\,\hat\sigma_{\mathrm{or}}/\sqrt{n}$ has asymptotic
coverage $1-\alpha$.
\end{proposition}

Proofs of both results are in Appendix~\ref{app:inference}. The proof of the
CLT proceeds by applying the delta method to $f(u,v)=u/(2v)$ after the
bivariate CLT for $(\bar U_n,\bar V_n)$; the cross-terms in the delta-method
expansion cancel when expressed in terms of the centred score $\psi_i$, leaving
the clean formula~\eqref{eq:clt}.

\subsection{Plug-in estimator and Neyman orthogonality}
\label{sec:plugin}
\noindent
When $m$ and $r$ are unknown, replace them with estimators $\hat m$ and
$\hat r$ to obtain
\begin{equation}
  \hat\tau:=\frac{\frac{1}{n}\sum_{i}(2p_i-1)(Y_i-\hat m(X_i))}
                 {2\,\frac{1}{n}\sum_{i}(p_i-\hat r(X_i))^{2}}.
  \label{eq:plugin}
\end{equation}
The denominator stability under nuisance estimation error is non-trivial and
is isolated as a separate lemma.

\begin{lemma}[Denominator stability]
\label{lem:denom}
If $p_i,\hat r(X_i)\in[0,1]$ a.s.\ and
$n^{-1}\sum_i(\hat r(X_i)-r(X_i))^{2}\xrightarrow{\;p\;} 0$, then
$n^{-1}\sum_i(p_i-\hat r(X_i))^{2}-n^{-1}\sum_i(p_i-r(X_i))^{2}\xrightarrow{\;p\;} 0$.
\end{lemma}

\begin{proposition}[Plug-in consistency]
\label{prop:plugin}
Under Assumptions~\ref{ass:mean}--\ref{ass:moments}, i.i.d.\ sampling,
$p_i,\hat r(X_i)\in[0,1]$ a.s., and
$n^{-1}\sum_i(\hat m(X_i)-m(X_i))^{2}\xrightarrow{\;p\;} 0$,
$n^{-1}\sum_i(\hat r(X_i)-r(X_i))^{2}\xrightarrow{\;p\;} 0$: $\;\hat\tau\xrightarrow{\;p\;}\tau$.
\end{proposition}

Proofs are in Appendix~\ref{app:plugin}.

For $\sqrt{n}$-normality of $\hat\tau$ with nuisances estimated at nonparametric
rates, the score~\eqref{eq:score} must be Neyman-orthogonal
\citep{chernozhukov2018}. Appendix~\ref{app:ortho} verifies that the
$r$-Gateaux derivative of $\mathbb{E}[\psi]$ is already zero (because $\mathbb{E}[a\mid X]=0$),
while the $m$-Gateaux derivative equals $-\mathbb{E}[(2r(X)-1)\delta_m(X)]$, which is
non-zero whenever $r(X)\not\equiv\frac12$. The score therefore fails Neyman
orthogonality through its $m$-direction.

Appendix~\ref{app:ortho} also identifies a natural Neyman-orthogonal
reformulation. Replacing $(2p-1)$ by $2(p-r(X))$ in the numerator gives the
score $\tilde\psi_i:=2(p_i-r(X_i))(Y_i-m(X_i)-\tau(p_i-r(X_i)))$, which has
both Gateaux derivatives equal to zero. The estimator defined by solving
$\frac{1}{n}\sum_i\tilde\psi_i(\hat\tau)=0$ is
\begin{equation}
  \hat\tau_{\mathrm{ort}} :=
  \frac{\frac{1}{n}\sum_i(p_i-\hat r(X_i))(Y_i-\hat m(X_i))}
       {\frac{1}{n}\sum_i(p_i-\hat r(X_i))^2},
  \label{eq:ort-estimator}
\end{equation}
which is a distinct estimator from~\eqref{eq:plugin}. When nuisances are
known, both estimators converge to $\tau$ and are asymptotically equivalent;
with estimated nuisances, $\hat\tau_{\mathrm{ort}}$ is the natural candidate for
DML-compatible inference, but establishing $\sqrt{n}$-normality of~\eqref{eq:ort-estimator}
under cross-fitting requires a separate proof that is beyond the scope of this paper.

\section{Robustness to Calibration Failure}
\label{sec:robustness}
\noindent
Suppose Assumption~\ref{ass:cal} is violated and
$\mathbb{E}[G\mid p,X]=g(p,X)=p+\eta(p,X)$ for a measurable calibration error
function $\eta$.

\begin{proposition}[Bias under calibration failure]
\label{prop:miscal}
Suppose Assumption~\ref{ass:mean} holds, $\mathbb{E}[G\mid p,X]=p+\eta(p,X)$, and
$V^{*}>0$. Then
\begin{equation}
  \operatorname{plim}\;\hat{\tau}_{\mathrm{or}}=\tau+B_{\mathrm{cal}},\qquad
  B_{\mathrm{cal}}:=\frac{\tau\,\mathbb{E}[(2p-1)\,\eta(p,X)]}{2V^{*}}.
  \label{eq:bias}
\end{equation}
\end{proposition}

The bias $B_{\mathrm{cal}}$ is proportional to $\tau$: no bias arises when
the true effect is zero, regardless of miscalibration. It is proportional to
the score-weighted mean of the calibration error; it vanishes whenever
$\mathbb{E}[(2p-1)\eta]=0$, which holds when the miscalibration is symmetric
in the sense of being orthogonal to the signed score.

\begin{proposition}[Sharp sensitivity bound]
\label{prop:sensitivity}
Suppose Assumption~\ref{ass:mean} holds and, for some $\delta\geq 0$,
$|\eta(p,X)|\leq\delta$ a.s. Then, over the class
$\mathcal{H}_\delta:=\{\eta:\,|\eta|\leq\delta\;\text{a.s.}\}$,
\begin{equation}
  \sup_{\eta\in\mathcal{H}_\delta}
  \left|\operatorname{plim}\;\hat{\tau}_{\mathrm{or}}-\tau\right|
  =|\tau|\cdot\frac{\delta\,\mathbb{E}[|2p-1|]}{2V^{*}}.
  \label{eq:sensitivity}
\end{equation}
The supremum is attained at $\eta^{*}(p,X)=\delta\,\operatorname{sgn}(2p-1)$.
\end{proposition}

The bound in~\eqref{eq:sensitivity} has a clean signal-to-noise
interpretation. The denominator $2V^{*}/\mathbb{E}[|z|]$ is an effective informativeness
measure of the score; larger $V^{*}$ means the score is more discriminating, and
the same calibration error $\delta$ produces proportionally less bias. As
$V^{*}\to 0$, the bound diverges, consistently with the identification failure of
Proposition~\ref{prop:failure}. Proofs are in Appendix~\ref{app:robust}.

\section{Monte Carlo Evidence}
\label{sec:mc}
\noindent
We report five sets of simulations, each tied directly to a theoretical result.
All experiments use $R=2{,}000$ replications with seeded random draws.
The baseline DGP has $X\in\mathbb{R}^3$ with independent standard normal entries,
$r(X)=\sigma(\beta_r^\top X)$ (logistic), and $m(X)=\beta_m^\top X$ (linear).
The score is drawn as $p\mid X\sim\mathrm{Beta}(r(X)\kappa,\,(1-r(X))\kappa)$
where $\kappa=(1-\sigma_u^2)/\sigma_u^2$, so that $\mathbb{E}[p\mid X]=r(X)$ and
$\operatorname{Var}(p\mid X)=\sigma_u^2\,r(X)(1-r(X))$ exactly, giving
$V^{*} = \sigma_u^2\,\mathbb{E}[r(X)(1-r(X))]$ exactly (not an approximation).
The outcome is $Y=m(X)+\tau G+\varepsilon$ with $G\sim\mathrm{Bernoulli}(p)$
and $\varepsilon\sim N(0,1)$.
The oracle estimator uses the true nuisance functions $m(X)$ and $r(X)$;
the plug-in estimator fits degree-2 polynomial ridge regressions without
cross-fitting; the orthogonal estimator uses 5-fold cross-fitting with the
same ridge models; and the hard-threshold estimator replaces $p$ with
$\mathbf{1}\{p>\frac12\}$.
Full replication code is provided in the online supplement.

\subsection{Finite-sample performance and oracle normality}
\label{sec:mc-baseline}

\begin{sloppypar}
\noindent
Table~\ref{tab:sim1} reports bias, standard deviation, RMSE, and empirical
coverage of nominal 95\% Wald intervals for the three main estimators at
$n\in\{500,1{,}000,5{,}000\}$ with $\tau=1$ and $\sigma_u=0.30$. The oracle
estimator is approximately unbiased throughout; the plug-in estimator exhibits
a persistent positive bias of roughly 0.12--0.17, attributable to in-sample
overfitting of the nuisance functions without cross-fitting; the orthogonal
estimator, which uses 5-fold cross-fitting, is nearly unbiased and achieves
coverage close to the nominal 0.95. RMSE shrinks at the $\sqrt{n}$ rate for
all three estimators.
\end{sloppypar}

Figure~\ref{fig:qq} shows normal QQ-plots of the standardised oracle estimates
$\sqrt{n}(\hat\tau_{\mathrm{or}}-\tau)/\hat\sigma_{\mathrm{or}}$ at each
sample size. The agreement with the $N(0,1)$ reference is excellent at
$n=1{,}000$ and $n=5{,}000$, confirming Theorem~\ref{thm:clt}.

\begin{table}[H]
\centering
\caption{Finite-sample performance under correct specification ($\tau=1$,
$\sigma_u=0.30$, $R=2{,}000$ replications). The orthogonal estimator uses
5-fold cross-fitting; the plug-in estimator does not, which accounts for its
persistent positive bias.}
\label{tab:sim1}
\begin{threeparttable}
\begin{tabular}{llrrrr}
\toprule
Estimator & $n$ & Bias & SD & RMSE & Coverage \\ \midrule
Oracle     & 500   & $+$0.004 & 0.501 & 0.501 & 0.952 \\
Plug-in    & 500   & $+$0.170 & 0.359 & 0.397 & 0.987 \\
Orthogonal & 500   & $-$0.014 & 0.329 & 0.329 & 0.955 \\ \midrule
Oracle     & 1,000 & $-$0.016 & 0.365 & 0.365 & 0.945 \\
Plug-in    & 1,000 & $+$0.150 & 0.252 & 0.294 & 0.983 \\
Orthogonal & 1,000 & $-$0.010 & 0.233 & 0.233 & 0.958 \\ \midrule
Oracle     & 5,000 & $-$0.005 & 0.162 & 0.163 & 0.946 \\
Plug-in    & 5,000 & $+$0.119 & 0.114 & 0.165 & 0.956 \\
Orthogonal & 5,000 & $-$0.005 & 0.110 & 0.110 & 0.946 \\
\bottomrule
\end{tabular}
\begin{tablenotes}\small
\item \textit{Notes}: Coverage is the empirical frequency of nominal 95\%
Wald intervals over $R=2{,}000$ replications.
\end{tablenotes}
\end{threeparttable}
\end{table}

\begin{figure}[H]
\centering
\includegraphics[width=\textwidth]{figures/fig1_qq.pdf}
\caption{Normal QQ-plots of standardised oracle estimates
$\sqrt{n}(\hat\tau_{\mathrm{or}}-\tau)/\hat\sigma_{\mathrm{or}}$
at $n\in\{500,1{,}000,5{,}000\}$. Confirming Theorem~\ref{thm:clt}.}
\label{fig:qq}
\end{figure}

\subsection{Approach to the identification boundary}
\label{sec:mc-boundary}

\noindent
Proposition~\ref{prop:failure} predicts that the estimator is not identified
when $V^{*}=0$ and that RMSE diverges as $V^{*}\to 0$.
Table~\ref{tab:sim2} traces this by decreasing score noise $\sigma_u$ at
$n=1{,}000$. As $V^{*}$ falls from $5.6\times 10^{-2}$ to $2.3\times 10^{-7}$,
RMSE grows by five orders of magnitude, while coverage remains close to
its nominal level throughout---the widening confidence intervals correctly
track the growing variance. Figure~\ref{fig:boundary} plots RMSE on a
log-log scale (left) and CI coverage (right); the empirical RMSE tracks
the theoretical $\mathrm{RMSE}\propto 1/V^{*}$ reference closely.

\begin{table}[H]
\centering
\caption{Approach to the identification boundary ($n=1{,}000$, $\tau=1$,
$R=2{,}000$ replications). Oracle estimator.}
\label{tab:sim2}
\begin{threeparttable}
\begin{tabular}{rrrrrr}
\toprule
$\sigma_u$ & True $V^{*}$ & Bias & SD & RMSE & Coverage \\ \midrule
0.500 & $5.63\times 10^{-2}$ & $-$0.001 & 0.164 & 0.164 & 0.957 \\
0.250 & $1.41\times 10^{-2}$ & $+$0.003 & 0.475 & 0.475 & 0.956 \\
0.100 & $2.25\times 10^{-3}$ & $+$0.044 & 2.578 & 2.578 & 0.953 \\
0.050 & $5.63\times 10^{-4}$ & $-$0.318 & 9.559 & 9.562 & 0.952 \\
0.010 & $2.25\times 10^{-5}$ & $+$6.363 & 242.8 & 242.8 & 0.946 \\
0.005 & $5.61\times 10^{-6}$ & $+$17.11 & 978.3 & 978.2 & 0.952 \\
0.001 & $2.25\times 10^{-7}$ & $-$158.3 & 24858 & 24853 & 0.947 \\
\bottomrule
\end{tabular}
\begin{tablenotes}\small
\item \textit{Notes}: $V^{*}=\mathbb{E}[(p-r(X))^2]$. SD, RMSE, and coverage are
computed on finite estimates only.
\end{tablenotes}
\end{threeparttable}
\end{table}

\begin{figure}[H]
\centering
\includegraphics[width=\textwidth]{figures/fig2_combined.pdf}
\caption{Identification boundary. Left: empirical RMSE (solid) and the
theoretical $1/V^{*}$ reference (dashed) on a log-log scale. Right:
CI coverage as $V^{*}\to 0$. Oracle estimator, $n=1{,}000$.
Confirms Proposition~\ref{prop:failure}.}
\label{fig:boundary}
\end{figure}

\subsection{Calibration failure and the sensitivity bound}
\label{sec:mc-calibration}

\noindent
Propositions~\ref{prop:miscal} and~\ref{prop:sensitivity} characterise bias
under miscalibration and show that the bound
$|\tau|\,\delta\,\mathbb{E}[|2p-1|]/(2V^{*})$ is sharp over $\mathcal{H}_\delta$.
Table~\ref{tab:sim3} and Figure~\ref{fig:bias} evaluate three calibration
error shapes at four values of $\delta$, with $n=2{,}000$ and $\tau=1$.
The \emph{worst-case} shape $\eta=\delta\,\mathrm{sgn}(2p-1)$ nearly attains
the sharp bound: tightness ratios range from 0.86 to 0.99, reflecting
Monte Carlo sampling variation. The \emph{symmetric} shape
$\eta=\delta\sin(\pi p)$ satisfies $\mathbb{E}[(2p-1)\eta]\approx 0$ and produces
near-zero theoretical and empirical bias regardless of $\delta$, confirming
that calibration errors orthogonal to the signed score leave the estimator
unbiased. The \emph{linear} shape lies between these extremes.

\begin{table}[H]
\centering
\caption{Calibration failure and sharp sensitivity bound ($n=2{,}000$,
$\tau=1$, $\sigma_u=0.30$, $R=2{,}000$ replications).}
\label{tab:sim3}
\begin{threeparttable}
\begin{tabular}{llrrrr}
\toprule
$\eta$ shape & $\delta$ & Emp.\ bias & Theo.\ bias & Sharp bound & Tightness \\
\midrule
Worst-case
  & 0.05 & $+$0.434 & $+$0.439 & 0.439 & 0.989 \\
  & 0.10 & $+$0.852 & $+$0.879 & 0.879 & 0.970 \\
  & 0.15 & $+$1.199 & $+$1.318 & 1.318 & 0.910 \\
  & 0.20 & $+$1.515 & $+$1.756 & 1.756 & 0.862 \\ \midrule
Linear
  & 0.05 & $+$0.232 & $+$0.226 & 0.439 & 0.528 \\
  & 0.10 & $+$0.423 & $+$0.451 & 0.878 & 0.482 \\
  & 0.15 & $+$0.602 & $+$0.676 & 1.316 & 0.457 \\
  & 0.20 & $+$0.774 & $+$0.903 & 1.757 & 0.440 \\ \midrule
Symmetric
  & 0.05 & $-$0.009 & $\approx$0 & 0.439 & 0.020 \\
  & 0.10 & $-$0.003 & $\approx$0 & 0.878 & 0.004 \\
  & 0.15 & $+$0.000 & $\approx$0 & 1.317 & 0.000 \\
  & 0.20 & $+$0.008 & $\approx$0 & 1.757 & 0.004 \\
\bottomrule
\end{tabular}
\begin{tablenotes}\small
\item \textit{Notes}: Worst-case: $\eta=\delta\,\mathrm{sgn}(2p-1)$.
Linear: $\eta=\delta(2p-1)$. Symmetric: $\eta=\delta\sin(\pi p)$.
Theoretical bias: $B_{\mathrm{cal}}=\tau\mathbb{E}[(2p-1)\eta]/(2V^{*})$.
Sharp bound: $|\tau|\,\delta\,\mathbb{E}[|2p-1|]/(2V^{*})$.
Tightness $=|$emp.\ bias$|/$sharp bound. Oracle estimator.
\end{tablenotes}
\end{threeparttable}
\end{table}

\begin{figure}[H]
\centering
\includegraphics[width=\textwidth]{figures/fig3_bias_bound.pdf}
\caption{Empirical bias (points) and sharp bound (dashed) as functions of
$\delta$ for three calibration error shapes. The worst-case shape nearly
attains the bound; the symmetric shape generates bias indistinguishable
from zero. Confirms Propositions~\ref{prop:miscal}
and~\ref{prop:sensitivity}.}
\label{fig:bias}
\end{figure}

\subsection{Attenuation by hard-threshold classification}
\label{sec:mc-threshold}

\noindent
Table~\ref{tab:sim4} and Figure~\ref{fig:attenuation} evaluate the attenuation
result in a DGP satisfying $r(X)=\frac12$ a.s.\ and the conditional symmetry
condition of Appendix~\ref{app:threshold}, at $\sigma_u\in\{0.10,0.20,0.30\}$
with $n=1{,}000$ and $\tau=1$. The oracle and orthogonal estimators are
centred on $\tau=1$ in every setting. The threshold estimator converges to
approximately $\hat\kappa\tau$ and attenuation worsens sharply as $\sigma_u$
decreases: at $\sigma_u=0.10$, the threshold estimate is approximately $0.08$
where the truth is $1$.

\begin{table}[H]
\centering
\caption{Hard-threshold versus moment estimators ($n=1{,}000$, $\tau=1$,
$R=2{,}000$ replications).}
\label{tab:sim4}
\begin{threeparttable}
\begin{tabular}{rllrrr}
\toprule
$\sigma_u$ & $\hat\kappa$ & Estimator & Mean $\hat\tau$ & Bias & RMSE \\
\midrule
0.10 & 0.100 & Oracle    & 1.008 & $+$0.008 & 0.889 \\
     &       & Plug-in   & 0.897 & $-$0.103 & 0.719 \\
     &       & Threshold & 0.085 & $-$0.915 & 0.919 \\ \midrule
0.20 & 0.172 & Oracle    & 1.005 & $+$0.005 & 0.379 \\
     &       & Plug-in   & 0.944 & $-$0.056 & 0.353 \\
     &       & Threshold & 0.163 & $-$0.837 & 0.840 \\ \midrule
0.30 & 0.252 & Oracle    & 1.005 & $+$0.005 & 0.239 \\
     &       & Plug-in   & 0.952 & $-$0.048 & 0.232 \\
     &       & Threshold & 0.249 & $-$0.751 & 0.754 \\
\bottomrule
\end{tabular}
\begin{tablenotes}\small
\item \textit{Notes}: $\hat\kappa=2\,\overline{|p-\frac12|}$ is the
empirical attenuation factor. Under the conditions of
Appendix~\ref{app:threshold}, $\operatorname{plim}\,\hat\tau_{\mathrm{ht}}=\kappa\tau$.
Plug-in: degree-2 polynomial ridge models without cross-fitting.
\end{tablenotes}
\end{threeparttable}
\end{table}

\begin{figure}[H]
\centering
\includegraphics[width=\textwidth]{figures/fig4_attenuation.pdf}
\caption{Sampling distributions of the oracle, plug-in, and hard-threshold
estimators at three levels of score dispersion $\sigma_u$. The threshold
estimator is centred well below $\tau=1$ in each panel; attenuation worsens
as $\sigma_u$ decreases. Confirms Appendix~\ref{app:threshold}.}
\label{fig:attenuation}
\end{figure}

\subsection{Heterogeneous effects and the variance-weighted estimand}
\label{sec:mc-hetero}

\noindent
\begin{sloppypar}
The Discussion in Section~\ref{sec:discussion} establishes that when the
structural effect varies with covariates, the moment estimator identifies the
variance-weighted average $\bar\tau=\mathbb{E}[\tau(X)\operatorname{Var}(p\mid X)]/\mathbb{E}[\operatorname{Var}(p\mid X)]$
rather than the simple mean $\mathbb{E}[\tau(X)]$. Table~\ref{tab:sim5} and
Figure~\ref{fig:weighted} confirm this in a DGP with
$\tau(X)=\tau_0+\tau_1 X_1$, $\tau_0=1$, $\tau_1=0.5$.
Design A holds $\operatorname{Var}(p\mid X)$ constant, so $\bar\tau=\tau_0=1.001$. Design B
introduces $X$-varying score variance $\operatorname{Var}(p\mid X)\propto e^{0.8X_1}$, which
upweights units with large positive $X_1$ and gives $\bar\tau=1.362\neq\mathbb{E}[\tau(X)]=1$.
In both designs, bias is negligible and shrinks towards zero as $n$ grows,
confirming that the oracle estimator correctly identifies the variance-weighted
estimand.
\end{sloppypar}

\begin{table}[H]
\centering
\caption{Heterogeneous effects: variance-weighted estimand recovery
($\tau(X)=\tau_0+\tau_1 X_1$, $\tau_0=1$, $\tau_1=0.5$,
$R=2{,}000$ replications).}
\label{tab:sim5}
\begin{threeparttable}
\begin{tabular}{llrrrr}
\toprule
Design & $n$ & $\bar\tau$ & Bias & RMSE & Coverage \\ \midrule
A: constant $\operatorname{Var}(p\mid X)$
  & 500   & 1.001 & $-$0.019 & 0.517 & 0.956 \\
  & 1,000 & 1.001 & $-$0.009 & 0.378 & 0.942 \\
  & 5,000 & 1.001 & $+$0.000 & 0.168 & 0.942 \\ \midrule
B: heterogeneous $\operatorname{Var}(p\mid X)$
  & 500   & 1.362 & $-$0.006 & 0.416 & 0.950 \\
  & 1,000 & 1.362 & $-$0.004 & 0.296 & 0.951 \\
  & 5,000 & 1.362 & $+$0.001 & 0.131 & 0.953 \\
\bottomrule
\end{tabular}
\begin{tablenotes}\small
\item \textit{Notes}: $\bar\tau=\mathbb{E}[\tau(X)\operatorname{Var}(p\mid X)]/\mathbb{E}[\operatorname{Var}(p\mid X)]$.
Design A: $\operatorname{Var}(p\mid X)$ constant; $\bar\tau=\tau_0=1$.
Design B: $\operatorname{Var}(p\mid X)=\sigma_u^2 e^{0.8X_1}$; $\bar\tau\neq\mathbb{E}[\tau(X)]=1$.
Oracle estimator with true $m(X)$ and $r(X)$.
\end{tablenotes}
\end{threeparttable}
\end{table}

\begin{figure}[H]
\centering
\includegraphics[width=.75\textwidth]{figures/fig5_weighted.pdf}
\caption{RMSE relative to the variance-weighted estimand $\bar\tau$ as a
function of sample size, for Design A (constant weights, $\bar\tau=\tau_0=1$)
and Design B (heterogeneous weights, $\bar\tau=1.362$). Both designs converge
to their respective $\bar\tau$, confirming the heterogeneous-effects discussion
in Section~\ref{sec:discussion}.}
\label{fig:weighted}
\end{figure}

\section{Discussion}
\label{sec:discussion}
\noindent
We have established point identification of a structural latent-group
coefficient $\tau$, a sharp characterisation of identification failure,
oracle inference and plug-in consistency results, and robustness
to calibration failure. This section discusses three further topics: the
attenuation induced by hard-threshold classification, the interpretation of
the estimand under heterogeneous effects, and three directions open for
subsequent research.

A common alternative to the moment estimator~\eqref{eq:oracle} is to threshold
the score at $p=\frac12$, form a binary indicator $\tilde G:=\mathbf{1}\{p>\frac12\}$,
and estimate the group gap as the difference in conditional means across the two
induced cells. Appendix~\ref{app:threshold} shows that this estimator converges
to $\kappa\tau$ with $\kappa=2\mathbb{E}[|p-\frac12|]\in(0,1)$ under mild conditions,
so the moment estimator strictly dominates whenever classification is imperfect.
The Monte Carlo evidence in Section~\ref{sec:mc-threshold} confirms that
attenuation can be severe: when score dispersion is low the threshold estimator
recovers less than ten percent of the true coefficient.

If the structural mean model admits a covariate-varying effect,
$\mathbb{E}[Y\mid G,p,X]=\mu(X)+\tau(X)G$, the same proof strategy shows that the
moment equation identifies the variance-weighted average
$\bar\tau:=\mathbb{E}[\tau(X)\operatorname{Var}(p\mid X)]/\mathbb{E}[\operatorname{Var}(p\mid X)]$, where the weight
$\operatorname{Var}(p\mid X)$ is the local informativeness of the score at covariate value
$X$. Under the constant-coefficient restriction this reduces to $\tau$.
The weight function has a natural interpretation: units whose score varies
substantially beyond what is predicted by covariates contribute more
information to the identification of the group effect, and the estimand
accordingly upweights their individual effects.

Three directions are natural for subsequent work. First, the orthogonal
estimator~\eqref{eq:ort-estimator}, whose score is Neyman-orthogonal, is
the natural candidate for $\sqrt{n}$-normality under nonparametric nuisance
estimation via cross-fitting; establishing this formally requires verifying
the DML remainder conditions of \citet{chernozhukov2018}, which is left to
subsequent work. Second, whether the oracle estimator achieves the
semiparametric efficiency bound for the model defined by
Assumptions~\ref{ass:mean}--\ref{ass:var} requires a tangent-space calculation
not undertaken here. Third, our sensitivity bounds are sharp over
$\mathcal{H}_\delta$, but tighter bounds may be achievable under shape
restrictions on $\eta(p,X)$.

\section{Conclusion}
\label{sec:conclusion}
\noindent
This paper has developed a framework for identifying and estimating a
structural group effect when the binary group indicator is latent but a
calibrated probability score is observed. Under a constant-coefficient
conditional mean model and the calibration condition $\mathbb{E}[G\mid p,X]=p$, the
structural coefficient $\tau$ is point-identified by a closed-form ratio of
observable moments, provided the score carries residual variation beyond
covariates. Identification fails precisely when this residual variation is
absent, and the failure is characterised constructively by an explicit family
of observationally equivalent models with arbitrary coefficients.

Several conclusions follow from the analysis. The identified coefficient is a
within-covariate-cell structural effect, distinct from the marginal group mean
gap; the two coincide if and only if the latent groups are covariate-balanced.
The oracle estimator is $\sqrt{n}$-consistent and asymptotically normal with a
closed-form sandwich variance, and the moment approach strictly dominates
hard-threshold classification whenever the score is imperfectly concentrated.
When calibration is imperfect, the bias admits an exact formula and is bounded
by a sharp sensitivity bound that scales inversely with the residual score
variance, consistently with the identification result.

The Monte Carlo evidence confirms each of these predictions quantitatively:
the oracle estimator is approximately unbiased and asymptotically normal, RMSE
diverges at the predicted rate as $V^{*}\to 0$, calibration errors produce bias
bounded by the sharp formula, and hard-threshold classification induces the
predicted attenuation factor $\kappa$.

Looking forward, the most important open direction is the formal verification
of $\sqrt{n}$-normality for the Neyman-orthogonal estimator under cross-fitting,
which would place the approach within the double machine learning framework of
\citet{chernozhukov2018} and open the path to inference with flexible nuisance
estimators. The framework developed here---centred on a calibrated probability
as a proxy for latent membership---has natural applications in distributional
analysis, fairness auditing, and any empirical setting where group indicators
are administratively missing but predictable from observed characteristics.

\clearpage
\bibliographystyle{elsarticle-harv}
\bibliography{references}