Yu Jeffrey Hu, Jeroen Rombouts, Ines Wilms
arXiv 23 Apr 2025 · Econometrics
arXiv:2504.16789 · PDF · DOI · OpenAlex · Extracted main text
Machine learning models are widely recognized for their strong performance in forecasting. To keep that performance in streaming data settings, they have to be monitored and frequently re-trained. This can be done with machine learning operations (MLOps) techniques under supervision of an MLOps engineer. However, in digital platform settings where the number of data streams is typically large and unstable, standard monitoring becomes either suboptimal or too labor intensive for the MLOps engineer. As a consequence, companies often fall back on very simple worse performing ML models without monitoring. We solve this problem by adopting a design science approach and introducing a new monitoring framework, the Machine Learning Monitoring Agent (MLMA), that is designed to work at scale for any ML model with reasonable labor cost. A key feature of our framework concerns test-based automated re-training based on a data-adaptive reference loss batch. The MLOps engineer is kept in the loop via key metrics and also acts, pro-actively or retrospectively, to maintain performance of the ML model in the production stage. We conduct a large-scale test at a last-mile delivery platform to empirically validate our monitoring framework.
appendix boundary found by appendix_command · 93% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Kreuzberger, D., Kühl, N., and Hirschl, S (2023) Machine learning operations (MLOps): Overview, definition, and architecture | 0.874 | 5 | 2 | 100% |
| 2 | Luo, L., Zhou, L., and Song, P. X.-K (2023) Real-time regression analysis of streaming clustered data with possible abnormal data batches | 0.843 | 4 | 3 | 75% |
| 3 | Testi, M., Ballabio, M., Frontoni, E., Iannello, G., Moccia, S., Sod… (2022) MLOps: A taxonomy and a methodology | 0.811 | 4 | 2 | 100% |
| 4 | Das, S. D. and Bala, P. K (2024) What drives MLOps adoption? An analysis using the TOE framework | 0.644 | 2 | 2 | 100% |
| 5 | Hu, Y. J., Rombouts, J., and Wilms, I (2024) Fast forecasting of unstable data streams for on-demand service platforms self | 0.644 | 2 | 2 | 100% |
| 6 | Ruf, P., Madan, M., Reich, C., and Ould-Abdeslam, D (2021) Demystifying MLOps and presenting a recipe for the selection of open-source tools | 0.644 | 2 | 2 | 100% |
| 7 | Chu, C.-S. J., Stinchcombe, M., and White, H (1996) Monitoring structural change | 0.511 | 2 | 2 | 50% |
| 8 | Ahmad, S., Lavin, A., Purdy, S., and Agha, Z (2017) Unsupervised real-time anomaly detection for streaming data | 0.405 | 1 | 1 | 100% |
| 9 | Balakrishnan, M., Ferreira, K. J., and Tong, J (2025) Human-algorithm collaboration with private information: Naïve advice-weighting behavior and mitigation | 0.405 | 1 | 1 | 100% |
| 10 | Benbya, H., Nan, N., Tanriverdi, H., and Yoo, Y (2020) Complexity and information systems research in the emerging digital world | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 53 scored citations.