EconBase
← All papers

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

Xinrui Ruan, Zhenyu Zhao, Waverly Wei, Yueshan Zhang, Zeyu Zheng, Sui Huang, Jingshen Wang

arXiv 29 Jun 2026 · Statistics — Methodology

arXiv:2606.29784 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Organizations thus often turn to collecting vast but noisy "silver" labels from crowdsourced workers or vendor annotators as proxies for gold labels. Because gold remains the evaluation target, naively aggregating noisy silver labels may introduce bias, and estimators built on sparsely observed gold labels may have high variance to resolve the model performance gaps that guide practical decisions. Model evaluation has become an ongoing operational practice rather than a one-time exercise, with evaluation rounds repeating across model versions, releases, and content domains. A natural question is whether the previous historical evaluation data can be used to improve each new round of evaluation. We introduce HERO (History Enhanced RObust model evaluation), a novel framework that uses historical data to suppress bias (improve reliability) and reduce variance (improve sensitivity) in model performance evaluation. HERO calibrates silver labelers' performance learned from historical gold annotations, and stabilizes the resulting estimator by anchoring it to covariate information measured with high precision in the historical data. HERO can be broadly applied across multiple common evaluation tasks, and remains valid when only a subset of historical labelers appears in the current round. We establish conditions under which the bias and variance reductions hold, showcase HERO's performance in simulation studies, and demonstrate its effectiveness on real-world model evaluation benchmarking datasets.

Citation extraction

30
references
41
in-text mentions
30
distinct cited
1
self-citations
5,627
main-text words

appendix boundary found by appendix_command · 52% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Dawid, A. P. and Skene, A. M (1979) Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm0.7373367%
2Chiang, Wei-Lin and Zheng, Lianmin and Sheng, Ying and Angelopoulos,… (2024) Chatbot arena: An open platform for evaluating llms by human preference0.6443267%
3Snow, Rion and O’connor, Brendan and Jurafsky, Dan and Ng, Andrew Y (2008) Cheap and fast–but is it good? evaluating non-expert annotations for natural language tasks0.58531100%
4Raykar, Vikas C and Yu, Shipeng and Zhao, Linda H and Valadez, Gerar… (2010) Learning from crowds.0.5112250%
5Whitehill, Jacob and Wu, Ting-fan and Bergsma, Jacob and Movellan, J… (2009) Whose vote should count more: Optimal integration of labels from labelers of unknown expertise0.5112250%
6Cohen, Jacob (1960) A Coefficient of Agreement for Nominal Scales0.40511100%
7Fleiss, Joseph L (1971) Measuring Nominal Scale Agreement Among Many Raters0.40511100%
8Raykar, Vikas C. and Yu, Shipeng and Zhao, Linda H. and Valadez, Ger… (2010) Learning From Crowds0.40511100%
9Snow, Rion and O'Connor, Brendan and Jurafsky, Daniel and Ng, Andrew Y (2008) Cheap and Fast–-But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks0.40511100%
10Warfield, Simon K. and Zou, Kelly H. and Wells, William M (2004) An Algorithm for the Validation of Image Segmentation0.40511100%

Showing the top 10 of 30 scored citations.