EconBase
← All papers

Partial Identification from LLM Prompts

Xiaohong Chen, Ashesh Rambachan, Elie Tamer

arXiv 13 Jun 2026 · Econometrics

arXiv:2606.15031 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Large language models are increasingly used as binary classifiers when the true label is latent. We study partial identification of the prevalence $θ= P(X^* = 1)$ from panels of LLM reports whose errors may be arbitrarily dependent given the truth. The design of replication determines the observable, and hence the identifying content: repeated prompts to one model yield a count, several named models a response vector, and both a response matrix. Cast as a two-component finite mixture, the problem makes the identification failure transparent: absent restrictions that separate the latent components, the prevalence $θ$ is completely unidentified, and weak stochastic-ordering restrictions (first-order dominance, monotone likelihood ratio, mean ordering) leave the identified set at $[0,1]$. Identifying power comes instead from externally calibrated scores and events, which discipline the mixture in the spirit of the misclassification and corrupted-data literature. We characterize the resulting bounds, establishing validity and sharpness, and give an exact account of the identifying information in the full score distribution beyond its mean. When named models are asked repeated versions of the same question, what identifies $θ$ is not the number of positive answers but which models agree across prompts -- a feature a vote count discards. An extension derives implied bounds on regression coefficients when $X^*$ is a regressor of interest that is not directly observed.

Citation extraction

13
references
16
in-text mentions
13
distinct cited
1
self-citations
11,751
main-text words

appendix boundary found by appendix_command · 99% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Hu, Yingyao (2008) Identification and Estimation of Nonlinear Models with Misclassification Error Using Instrumental Variables: A General Solution0.64422100%
2Molinari, Francesca (2008) Partial Identification of Probability Distributions with Misclassified Data0.64422100%
3Henry, Marc and Kitamura, Yuichi and Salanié, Bernard (2014) Partial Identification of Finite Mixtures in Econometric Models0.5112250%
4Bollinger, Christopher R (1996) Bounding Mean Regressions When a Binary Regressor Is Mismeasured0.40511100%
5Cheng, Zelei and Wu, Xian and Yu, Jiahao and Han, Shuo and Cai, Xin-… (2024) Soft-Label Integration for Robust Toxicity Classification0.40511100%
6Dawid, A. P. and Skene, A. M (1979) Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm0.40511100%
7Horowitz, Joel L. and Manski, Charles F (1995) Identification and Robustness with Contaminated and Corrupted Data0.40511100%
8Hovsepian, Karen and Liu, Di and Murugesan, San (2024) Label with Confidence: Effective Confidence Calibration and Ensembles in LLM-Powered Classification0.40511100%
9Mahajan, Aprajit (2006) Identification and Estimation of Regression Models with Misclassification0.40511100%
10Linder, Fridolin and Leeper, Thomas J. and Haimovich, Daniel and Tax… (2026) Unbiased Prevalence Estimation with Multicalibrated LLMs0.40511100%

Showing the top 10 of 13 scored citations.