Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
41,102 characters · 16 sections · 10 citation commands
Identification of Latent Group Effects under Conditional Calibration
A pervasive challenge in empirical work is the measurement of outcome differences between groups when group membership is not directly observed. Poverty status, immigration status, informal employment, fuel insecurity, and latent health conditions are leading examples. In such settings the analyst typically has access to a probability score $p_i\in[0,1]$ encoding belief that unit $i$ belongs to the group of interest, but never observes the binary indicator $G_i\in\{0,1\}$ itself.
The central question we address is: under what conditions, and by what formula, can a structural group effect be identified from the joint law of observables $(Y,X,p)$ when $G$ is never observed? We give a precise answer organised around three claims. First, a structural coefficient $\tau$ is point-identified under mild conditions. Second, identification fails in a characterisable and sharp way when exactly one of those conditions is violated. Third, the identified object is distinct from the marginal group mean gap in a way that can be made fully explicit.
The paper makes four contributions. The first is an identification result. Under a constant-coefficient structural mean model and the conditional calibration condition $\mathbb{E}[G\mid p,X]=p$, we prove that $\tau$ is identified by a weighted moment equation whose denominator $V^{*}=\mathbb{E}[(p-r(X))^{2}]$ is the residual variance of the score after partialling on $X$. The formula is in closed form and admits a transparent interpretation: it is formally analogous to an instrumental-variables estimand in which the score residual $a=p-r(X)$ plays the role of an instrument for the latent deviation $G-r(X)$; the calibration condition supplies the first-stage relevance and the mean-independence condition in the structural model supplies the exclusion restriction.
The second contribution is an exact characterisation of identification failure. We prove that $\tau$ is not identified if and only if $V^{*}=0$, i.e.\ the score is a deterministic function of $X$, and we make this concrete by constructing an explicit continuum of observationally equivalent models indexed by arbitrary values $\tau'\in\mathbb{R}$.
The third contribution is a clean separation between the identified structural coefficient and the marginal latent mean gap $\Delta_{\mathrm{marg}}=\mathbb{E}[Y\mid G=1]-\mathbb{E}[Y\mid G=0]$. We decompose $\Delta_{\mathrm{marg}}=\tau+C$ where the compositional term $C$ is not identified from $(Y,X,p)$ alone, and we give a necessary and sufficient condition for $C=0$.
The fourth contribution is oracle inference and robustness. We establish $\sqrt{n}$-asymptotic normality of the oracle estimator with an explicit sandwich variance, compute the exact probability limit under calibration failure, and derive a sensitivity bound that is sharp over the class of all calibration error functions bounded uniformly by $\delta$.
Our paper sits at the intersection of four strands of the literature. Within the misclassification literature, lewbel2007 showed that average treatment effects are attenuated under misclassification of a binary regressor and proposed corrections; mahajan2006 obtained identification using an instrumental variable; kasahara2022 extended this to the endogenous case. Our setting is complementary: instead of observing a noisy binary label, the analyst observes a calibrated probability for $G$, which changes both the identification argument and the identified object.
The proxy variable and measurement error literature hu2008,schennach2016 establishes nonparametric identification of full latent-variable distributions via rank conditions on integral operators. Our setting is more restrictive---we target only the scalar $\tau$---but our assumptions are correspondingly weaker and the identification formula is closed-form. The structure of our moment equation parallels the partially linear model robinson1988 and semiparametric IV newey1990; the key departure is that the “instrument” is not an observable variable but a residual derived from a calibrated probability. Finally, the algorithmic fairness literature kallus2022,chen2019 has studied disparity estimation with unobserved protected attributes under calibration-type assumptions; our contribution to that context is a formal identification-theoretic treatment with a closed-form formula and exact failure and sensitivity characterisations.
The remainder of the paper is organised as follows. Section (ref) sets up the model. Section (ref) proves identification and characterises failure. Section (ref) distinguishes the structural coefficient from the marginal gap. Section (ref) covers oracle inference and the plug-in estimator, including Neyman orthogonality. Section (ref) develops robustness to calibration failure. Section (ref) presents Monte Carlo evidence. Section (ref) discusses extensions and open directions. Section (ref) concludes. The Appendix contains all proofs.
Let $(\Omega,\mathcal{F},\mathbb{P})$ be a probability space. We observe i.i.d.\ draws $(Y_i,X_i,p_i)\in\mathbb{R}\times\mathcal{X}\times[0,1]$ for $i=1,\ldots,n$, from the joint distribution $P_{Y,X,p}$. There exists on the same probability space an unobserved binary variable $G_i\in\{0,1\}$; $G_i=1$ denotes membership in the latent group of interest. The measurable covariate space $\mathcal{X}$ is arbitrary.
We write $m(x):=\mathbb{E}[Y\mid X=x]$, $r(x):=\mathbb{E}[p\mid X=x]$, and $\pi(x):=\mathbb{E}[G\mid X=x]$ for the conditional mean functions. The derived quantities
satisfy $\mathbb{E}[R\mid X]=\mathbb{E}[a\mid X]=0$. The residual score variance
measures the variation in $p$ not explained by $X$; it is the key quantity governing identification.
Assumption (ref) has two components. The effect of latent membership on the conditional mean of $Y$ is constant in $X$. Additionally, the score $p$ is mean-independent of $Y$ once $(G,X)$ are known: conditional on true membership, the analyst's probability score conveys no further information about the expected outcome.
Assumption (ref) is the sole formal link between the latent indicator $G$ and the observed score $p$. It is a calibration condition: $p$ need not equal the propensity score $\mathbb{P}(G=1\mid X)$ but must be an unbiased predictor of $G$ given all observed information $(p,X)$. Scores arising from area-level prevalence rates, classifier outputs, or model-based predictions all satisfy this condition when they are well-calibrated.
Assumption (ref) implies square-integrability of all relevant quantities and is used in Sections (ref)--(ref). Asymptotic normality in Section (ref) uses the full fourth-moment condition; identification requires only second moments.
The following two lemmas are the structural backbone of the paper.
Proofs are in Appendix (ref). The content of Lemma (ref) is structural: the entire predictable part of the outcome residual $R$ is driven by the deviation of true group membership from its score-implied expectation, $G-r(X)$.
The proof of Theorem (ref) is in Appendix (ref). The identification formula (ref) has a transparent algebraic structure. The numerator $\mathbb{E}[zR]$ is the covariance between the signed score $z=2p-1$ and the outcome residual $R$, after partialling both on $X$. The denominator $2V^{*}=2\mathbb{E}[a^{2}]$ is twice the residual variance of the score. The ratio is therefore the slope of the regression of $R$ on $z$ in the covariate-partialled data, which by Lemma (ref) in the Appendix equals the slope of the regression of $R$ on the latent deviation $G-r(X)$. This is formally analogous to an IV estimand in which $a=p-r(X)$ acts as an instrument for $G-r(X)$: the calibration condition (Assumption (ref)) supplies the first-stage relevance, and the mean-independence condition in Assumption (ref) supplies the exclusion restriction.
We now show that Assumption (ref) is not merely a regularity condition but the exact boundary of identification.
Part (b) is the substantive non-identification claim. The explicit construction is in Appendix (ref). It defines a new model $\mathcal{M}'$ that retains the observable distribution of $(Y,X,p)$, sets $G':=\mathbf{1}\{U\leq p\}$ for an independent $U\sim\mathrm{Uniform}(0,1)$, and postulates the structural equation with coefficient $\tau'$ and $\mu'(X):=m(X)-\tau'r(X)$. The postulate is consistent because the implied observable regression of $Y$ on $X$ equals $m(X)$ for every $\tau'\in\mathbb{R}$, so $\mathcal{M}'$ satisfies both assumptions while being observationally indistinguishable from the original model.
A natural question is whether $\tau$ equals the marginal latent mean gap $\Delta_{\mathrm{marg}}:=\mathbb{E}[Y\mid G=1]-\mathbb{E}[Y\mid G=0]$. Under Assumption (ref), a direct calculation gives
The term $C$ captures differences in covariate composition across latent groups. It depends on the latent conditional distributions $\mathbb{P}_{X\mid G=g}$, which are not identified from $(Y,X,p)$ under our assumptions: many latent joint distributions of $(G,X)$ generate the same observable distribution.
The proof is immediate from (ref). The practical import is that $\tau$ identifies the within-covariate-cell group effect, while $\Delta_{\mathrm{marg}}$ conflates this with the compositional term $C$. Recovering $\Delta_{\mathrm{marg}}$ requires identifying $C$, which in turn requires knowledge of the latent-group covariate distributions---not available under our assumptions without further restrictions.
Suppose for this subsection that $m$ and $r$ are known. Define the oracle estimator
and the score evaluated at the true parameter,
By Theorem (ref), $\mathbb{E}[\psi_i]=0$.
The variance $\sigma^{2}_{\mathrm{or}}=J^{-2}\mathbb{E}[\psi_i^{2}]$ has the standard sandwich form with Jacobian $J=2V^{*}$. The identification condition $V^{*}>0$ is precisely the condition that $J\neq 0$, i.e., that the moment equation is locally informative about $\tau$ in a neighbourhood of the truth.
Proofs of both results are in Appendix (ref). The proof of the CLT proceeds by applying the delta method to $f(u,v)=u/(2v)$ after the bivariate CLT for $(\bar U_n,\bar V_n)$; the cross-terms in the delta-method expansion cancel when expressed in terms of the centred score $\psi_i$, leaving the clean formula (ref).
When $m$ and $r$ are unknown, replace them with estimators $\hat m$ and $\hat r$ to obtain
The denominator stability under nuisance estimation error is non-trivial and is isolated as a separate lemma.
Proofs are in Appendix (ref).
For $\sqrt{n}$-normality of $\hat\tau$ with nuisances estimated at nonparametric rates, the score (ref) must be Neyman-orthogonal chernozhukov2018. Appendix (ref) verifies that the $r$-Gateaux derivative of $\mathbb{E}[\psi]$ is already zero (because $\mathbb{E}[a\mid X]=0$), while the $m$-Gateaux derivative equals $-\mathbb{E}[(2r(X)-1)\delta_m(X)]$, which is non-zero whenever $r(X)\not\equiv\frac12$. The score therefore fails Neyman orthogonality through its $m$-direction.
Appendix (ref) also identifies a natural Neyman-orthogonal reformulation. Replacing $(2p-1)$ by $2(p-r(X))$ in the numerator gives the score $\tilde\psi_i:=2(p_i-r(X_i))(Y_i-m(X_i)-\tau(p_i-r(X_i)))$, which has both Gateaux derivatives equal to zero. The estimator defined by solving $\frac{1}{n}\sum_i\tilde\psi_i(\hat\tau)=0$ is
which is a distinct estimator from (ref). When nuisances are known, both estimators converge to $\tau$ and are asymptotically equivalent; with estimated nuisances, $\hat\tau_{\mathrm{ort}}$ is the natural candidate for DML-compatible inference, but establishing $\sqrt{n}$-normality of (ref) under cross-fitting requires a separate proof that is beyond the scope of this paper.
Suppose Assumption (ref) is violated and $\mathbb{E}[G\mid p,X]=g(p,X)=p+\eta(p,X)$ for a measurable calibration error function $\eta$.
The bias $B_{\mathrm{cal}}$ is proportional to $\tau$: no bias arises when the true effect is zero, regardless of miscalibration. It is proportional to the score-weighted mean of the calibration error; it vanishes whenever $\mathbb{E}[(2p-1)\eta]=0$, which holds when the miscalibration is symmetric in the sense of being orthogonal to the signed score.
The bound in (ref) has a clean signal-to-noise interpretation. The denominator $2V^{*}/\mathbb{E}[|z|]$ is an effective informativeness measure of the score; larger $V^{*}$ means the score is more discriminating, and the same calibration error $\delta$ produces proportionally less bias. As $V^{*}\to 0$, the bound diverges, consistently with the identification failure of Proposition (ref). Proofs are in Appendix (ref).
We report five sets of simulations, each tied directly to a theoretical result. All experiments use $R=2{,}000$ replications with seeded random draws. The baseline DGP has $X\in\mathbb{R}^3$ with independent standard normal entries, $r(X)=\sigma(\beta_r^\top X)$ (logistic), and $m(X)=\beta_m^\top X$ (linear). The score is drawn as $p\mid X\sim\mathrm{Beta}(r(X)\kappa,\,(1-r(X))\kappa)$ where $\kappa=(1-\sigma_u^2)/\sigma_u^2$, so that $\mathbb{E}[p\mid X]=r(X)$ and $\operatorname{Var}(p\mid X)=\sigma_u^2\,r(X)(1-r(X))$ exactly, giving $V^{*} = \sigma_u^2\,\mathbb{E}[r(X)(1-r(X))]$ exactly (not an approximation). The outcome is $Y=m(X)+\tau G+\varepsilon$ with $G\sim\mathrm{Bernoulli}(p)$ and $\varepsilon\sim N(0,1)$. The oracle estimator uses the true nuisance functions $m(X)$ and $r(X)$; the plug-in estimator fits degree-2 polynomial ridge regressions without cross-fitting; the orthogonal estimator uses 5-fold cross-fitting with the same ridge models; and the hard-threshold estimator replaces $p$ with $\mathbf{1}\{p>\frac12\}$. Full replication code is provided in the online supplement.
Figure (ref) shows normal QQ-plots of the standardised oracle estimates $\sqrt{n}(\hat\tau_{\mathrm{or}}-\tau)/\hat\sigma_{\mathrm{or}}$ at each sample size. The agreement with the $N(0,1)$ reference is excellent at $n=1{,}000$ and $n=5{,}000$, confirming Theorem (ref).
Proposition (ref) predicts that the estimator is not identified when $V^{*}=0$ and that RMSE diverges as $V^{*}\to 0$. Table (ref) traces this by decreasing score noise $\sigma_u$ at $n=1{,}000$. As $V^{*}$ falls from $5.6\times 10^{-2}$ to $2.3\times 10^{-7}$, RMSE grows by five orders of magnitude, while coverage remains close to its nominal level throughout---the widening confidence intervals correctly track the growing variance. Figure (ref) plots RMSE on a log-log scale (left) and CI coverage (right); the empirical RMSE tracks the theoretical $\mathrm{RMSE}\propto 1/V^{*}$ reference closely.
Propositions (ref) and (ref) characterise bias under miscalibration and show that the bound $|\tau|\,\delta\,\mathbb{E}[|2p-1|]/(2V^{*})$ is sharp over $\mathcal{H}_\delta$. Table (ref) and Figure (ref) evaluate three calibration error shapes at four values of $\delta$, with $n=2{,}000$ and $\tau=1$. The worst-case shape $\eta=\delta\,\mathrm{sgn}(2p-1)$ nearly attains the sharp bound: tightness ratios range from 0.86 to 0.99, reflecting Monte Carlo sampling variation. The symmetric shape $\eta=\delta\sin(\pi p)$ satisfies $\mathbb{E}[(2p-1)\eta]\approx 0$ and produces near-zero theoretical and empirical bias regardless of $\delta$, confirming that calibration errors orthogonal to the signed score leave the estimator unbiased. The linear shape lies between these extremes.
Table (ref) and Figure (ref) evaluate the attenuation result in a DGP satisfying $r(X)=\frac12$ a.s.\ and the conditional symmetry condition of Appendix (ref), at $\sigma_u\in\{0.10,0.20,0.30\}$ with $n=1{,}000$ and $\tau=1$. The oracle and orthogonal estimators are centred on $\tau=1$ in every setting. The threshold estimator converges to approximately $\hat\kappa\tau$ and attenuation worsens sharply as $\sigma_u$ decreases: at $\sigma_u=0.10$, the threshold estimate is approximately $0.08$ where the truth is $1$.
We have established point identification of a structural latent-group coefficient $\tau$, a sharp characterisation of identification failure, oracle inference and plug-in consistency results, and robustness to calibration failure. This section discusses three further topics: the attenuation induced by hard-threshold classification, the interpretation of the estimand under heterogeneous effects, and three directions open for subsequent research.
A common alternative to the moment estimator (ref) is to threshold the score at $p=\frac12$, form a binary indicator $\tilde G:=\mathbf{1}\{p>\frac12\}$, and estimate the group gap as the difference in conditional means across the two induced cells. Appendix (ref) shows that this estimator converges to $\kappa\tau$ with $\kappa=2\mathbb{E}[|p-\frac12|]\in(0,1)$ under mild conditions, so the moment estimator strictly dominates whenever classification is imperfect. The Monte Carlo evidence in Section (ref) confirms that attenuation can be severe: when score dispersion is low the threshold estimator recovers less than ten percent of the true coefficient.
If the structural mean model admits a covariate-varying effect, $\mathbb{E}[Y\mid G,p,X]=\mu(X)+\tau(X)G$, the same proof strategy shows that the moment equation identifies the variance-weighted average $\bar\tau:=\mathbb{E}[\tau(X)\operatorname{Var}(p\mid X)]/\mathbb{E}[\operatorname{Var}(p\mid X)]$, where the weight $\operatorname{Var}(p\mid X)$ is the local informativeness of the score at covariate value $X$. Under the constant-coefficient restriction this reduces to $\tau$. The weight function has a natural interpretation: units whose score varies substantially beyond what is predicted by covariates contribute more information to the identification of the group effect, and the estimand accordingly upweights their individual effects.
Three directions are natural for subsequent work. First, the orthogonal estimator (ref), whose score is Neyman-orthogonal, is the natural candidate for $\sqrt{n}$-normality under nonparametric nuisance estimation via cross-fitting; establishing this formally requires verifying the DML remainder conditions of chernozhukov2018, which is left to subsequent work. Second, whether the oracle estimator achieves the semiparametric efficiency bound for the model defined by Assumptions (ref)--(ref) requires a tangent-space calculation not undertaken here. Third, our sensitivity bounds are sharp over $\mathcal{H}_\delta$, but tighter bounds may be achievable under shape restrictions on $\eta(p,X)$.
This paper has developed a framework for identifying and estimating a structural group effect when the binary group indicator is latent but a calibrated probability score is observed. Under a constant-coefficient conditional mean model and the calibration condition $\mathbb{E}[G\mid p,X]=p$, the structural coefficient $\tau$ is point-identified by a closed-form ratio of observable moments, provided the score carries residual variation beyond covariates. Identification fails precisely when this residual variation is absent, and the failure is characterised constructively by an explicit family of observationally equivalent models with arbitrary coefficients.
Several conclusions follow from the analysis. The identified coefficient is a within-covariate-cell structural effect, distinct from the marginal group mean gap; the two coincide if and only if the latent groups are covariate-balanced. The oracle estimator is $\sqrt{n}$-consistent and asymptotically normal with a closed-form sandwich variance, and the moment approach strictly dominates hard-threshold classification whenever the score is imperfectly concentrated. When calibration is imperfect, the bias admits an exact formula and is bounded by a sharp sensitivity bound that scales inversely with the residual score variance, consistently with the identification result.
The Monte Carlo evidence confirms each of these predictions quantitatively: the oracle estimator is approximately unbiased and asymptotically normal, RMSE diverges at the predicted rate as $V^{*}\to 0$, calibration errors produce bias bounded by the sharp formula, and hard-threshold classification induces the predicted attenuation factor $\kappa$.
Looking forward, the most important open direction is the formal verification of $\sqrt{n}$-normality for the Neyman-orthogonal estimator under cross-fitting, which would place the approach within the double machine learning framework of chernozhukov2018 and open the path to inference with flexible nuisance estimators. The framework developed here---centred on a calibrated probability as a proxy for latent membership---has natural applications in distributional analysis, fairness auditing, and any empirical setting where group indicators are administratively missing but predictable from observed characteristics.