EconBase
← All papers

Some models are useful, but for how long?: A decision theoretic approach to choosing when to refit large-scale prediction models

Kentaro Hoffman, Stephen Salerno, Jeff Leek, Tyler McCormick

arXiv 22 May 2024 · Statistics — Methodology

arXiv:2405.13926 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Large-scale prediction models using tools from artificial intelligence (AI) or machine learning (ML) are increasingly common across a variety of industries and scientific domains. Despite their effectiveness, training AI and ML tools at scale can cost tens or hundreds of thousands of dollars (or more); and even after a model is trained, substantial resources must be invested to keep models up-to-date. This paper presents a decision-theoretic framework for deciding when to refit an AI/ML model when the goal is to perform unbiased statistical inference using partially AI/ML-generated data. Drawing on portfolio optimization theory, we treat the decision of {\it recalibrating} a model or statistical inference versus {\it refitting} the model as a choice between “investing” in one of two “assets.” One asset, recalibrating the model based on another model, is quick and relatively inexpensive but bears uncertainty from sampling and may not be robust to model drift. The other asset, {\it refitting} the model, is costly but removes the drift concern (though not statistical uncertainty from sampling). We present a framework for balancing these two potential investments while preserving statistical validity. We evaluate the framework using simulation and data on electricity usage and predicting flu trends.

Citation extraction

56
references
65
in-text mentions
56
distinct cited
4
self-citations
8,372
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Lazer, D., R. Kennedy, G. King, and A. Vespignani (2014) Google flu trends still appears sick: An evaluation of the 2013-2014 flu season0.64441100%
2Angelopoulos, A. N., S. Bates, C. Fannjiang, M. I. Jordan, and T. Zr… (2023) Prediction-powered inference0.64422100%
3Hoffman, K., S. Salerno, A. Afiaz, J. T. Leek, and T. H. McCormick (2024) Do we really even need data? self0.64422100%
4Angelopoulos, A. N., J. C. Duchi, and T. Zrnic (2024) PPI++: Efficient prediction-powered inference0.58531100%
5Fan, S., A. Visokay, K. Hoffman, S. Salerno, L. Liu, J. T. Leek, and… (2024) From narratives to numbers: Valid inference using language model predictions from verbal autopsy narratives self0.51121100%
6Bonhomme, S. and M. Weidner (2022) Minimizing sensitivity to model misspecification0.51121100%
7Wooldridge, J. M (2010) Econometric Analysis of Cross Section and Panel Data0.40511100%
8Sumi, S. M (2021) Automated prediction of voter's party affiliation using ai0.40511100%
9Quang, P. X (1985) Robust sequential testing0.40511100%
10White, H (1982) Maximum likelihood estimation of misspecified models0.40511100%

Showing the top 10 of 56 scored citations.