Xinrui Ruan, Zhenyu Zhao, Waverly Wei, Yueshan Zhang, Zeyu Zheng, Sui Huang, Jingshen Wang
arXiv 29 Jun 2026 · Statistics — Methodology
arXiv:2606.29784 · PDF · DOI · OpenAlex · Extracted main text
Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Organizations thus often turn to collecting vast but noisy "silver" labels from crowdsourced workers or vendor annotators as proxies for gold labels. Because gold remains the evaluation target, naively aggregating noisy silver labels may introduce bias, and estimators built on sparsely observed gold labels may have high variance to resolve the model performance gaps that guide practical decisions. Model evaluation has become an ongoing operational practice rather than a one-time exercise, with evaluation rounds repeating across model versions, releases, and content domains. A natural question is whether the previous historical evaluation data can be used to improve each new round of evaluation. We introduce HERO (History Enhanced RObust model evaluation), a novel framework that uses historical data to suppress bias (improve reliability) and reduce variance (improve sensitivity) in model performance evaluation. HERO calibrates silver labelers' performance learned from historical gold annotations, and stabilizes the resulting estimator by anchoring it to covariate information measured with high precision in the historical data. HERO can be broadly applied across multiple common evaluation tasks, and remains valid when only a subset of historical labelers appears in the current round. We establish conditions under which the bias and variance reductions hold, showcase HERO's performance in simulation studies, and demonstrate its effectiveness on real-world model evaluation benchmarking datasets.
appendix boundary found by appendix_command · 52% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Dawid, A. P. and Skene, A. M (1979) Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm | 0.737 | 3 | 3 | 67% |
| 2 | Chiang, Wei-Lin and Zheng, Lianmin and Sheng, Ying and Angelopoulos,… (2024) Chatbot arena: An open platform for evaluating llms by human preference | 0.644 | 3 | 2 | 67% |
| 3 | Snow, Rion and O’connor, Brendan and Jurafsky, Dan and Ng, Andrew Y (2008) Cheap and fast–but is it good? evaluating non-expert annotations for natural language tasks | 0.585 | 3 | 1 | 100% |
| 4 | Raykar, Vikas C and Yu, Shipeng and Zhao, Linda H and Valadez, Gerar… (2010) Learning from crowds. | 0.511 | 2 | 2 | 50% |
| 5 | Whitehill, Jacob and Wu, Ting-fan and Bergsma, Jacob and Movellan, J… (2009) Whose vote should count more: Optimal integration of labels from labelers of unknown expertise | 0.511 | 2 | 2 | 50% |
| 6 | Cohen, Jacob (1960) A Coefficient of Agreement for Nominal Scales | 0.405 | 1 | 1 | 100% |
| 7 | Fleiss, Joseph L (1971) Measuring Nominal Scale Agreement Among Many Raters | 0.405 | 1 | 1 | 100% |
| 8 | Raykar, Vikas C. and Yu, Shipeng and Zhao, Linda H. and Valadez, Ger… (2010) Learning From Crowds | 0.405 | 1 | 1 | 100% |
| 9 | Snow, Rion and O'Connor, Brendan and Jurafsky, Daniel and Ng, Andrew Y (2008) Cheap and Fast–-But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks | 0.405 | 1 | 1 | 100% |
| 10 | Warfield, Simon K. and Zou, Kelly H. and Wells, William M (2004) An Algorithm for the Validation of Image Segmentation | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 30 scored citations.