EconBase
← All papers

Identification of Latent Group Effects under Conditional Calibration

Marcell T. Kurbucz

arXiv 9 Apr 2026 · Statistics — Methodology

arXiv:2604.08798 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We study identification of a structural group effect when the group indicator $G\in{0,1}$ is unobserved but the analyst observes a calibrated probability score $p\in[0,1]$ satisfying $\mathbb{E}[G|p,X]=p$. Under a constant-coefficient structural mean model, the latent-group coefficient $τ$ is point-identified from the joint law of observables $(Y,X,p)$ by a simple ratio of weighted moments: the covariance of the signed score $2p-1$ with the covariate-partialled outcome, divided by twice the residual variance of the score after conditioning on covariates. Identification fails if and only if the score is a deterministic function of $X$; we establish this by constructing an explicit continuum of observationally equivalent models indexed by arbitrary values of $τ$. The identified coefficient differs from the marginal latent mean gap by a compositional term that is unidentified without further assumptions; we give a necessary and sufficient condition for the two to coincide. The oracle estimator is $\sqrt{n}$-consistent and asymptotically normal with a closed-form sandwich variance. Under calibration error bounded uniformly by $δ$, the bias is bounded by $|τ|\,\mathbb{E}[|2p-1|]\,δ\,(2V^*)^{-1}$, a bound that is sharp over all calibration error functions of that magnitude. Hard-threshold classification at $p=1/2$ attenuates the estimated gap by a factor strictly less than one. Monte Carlo experiments confirm the asymptotic theory, trace the divergence of RMSE as $V^*\to 0$, illustrate the attenuation bias of hard-threshold classification, and verify identification of the variance-weighted estimand under heterogeneous effects.

Citation extraction

10
references
13
in-text mentions
10
distinct cited
0
self-citations
5,817
main-text words

appendix boundary found by appendix_command · 68% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters0.8434475%
2Chen, I.Y., Johansson, F.D., Sontag, D (2018) Why is my classifier discriminatory?0.40511100%
3Hu, Y., Schennach, S.M (2008) Instrumental variable treatment of nonclassical measurement error models0.40511100%
4Kallus, N., Mao, X., Zhou, A (2022) Assessing algorithmic fairness with unobserved protected class using data combination0.40511100%
5Kasahara, H., Shimotsu, K (2022) Identification of regression models with a misclassified and endogenous binary regressor0.40511100%
6Lewbel, A (2007) Estimation of average treatment effects with misclassification0.40511100%
7Mahajan, A (2006) Identification and estimation of regression models with misclassification0.40511100%
8Newey, W.K (1990) Efficient instrumental variables estimation of nonlinear models0.40511100%
9Robinson, P.M (1988) Root-$N$-consistent semiparametric regression0.40511100%
10Schennach, S.M (2016) Recent advances in the measurement error literature0.40511100%

Showing the top 10 of 10 scored citations.