EconBase
← All papers

The Noise Is the Signal: Correlated Sampling Error Is Rank-Informative for Proxy Metric Selection

Sandro Provenzano

arXiv 6 Oct 2026 · Econometrics

arXiv:2610.08194 · PDF · Extracted main text

Abstract

North-star metrics such as customer lifetime value are often too slow and noisy to decide a short A/B test. Teams therefore rely on a proxy metric, commonly chosen by how closely its effects tracked the north star's across past experiments. Validating that choice, or any method for making it, is hard: the only benchmark is the noisy north star, and the number of available past experiments is limited. In addition, proxy and north-star effects are estimated on the same customers, so their sampling errors are correlated. Recent work at major experimentation platforms removes this shared error as contamination, improving estimates of the true-effect covariance. Choosing a proxy, however, is a ranking problem, and a better estimate need not give a better ranking. We measure agreement free of shared error by estimating the two effects on disjoint random halves of each experiment's customers. In an archive of 262 experiments and 69 candidate proxies, the shared error ranks the candidates in a similar order to this agreement (Spearman correlation 0.65): it carries information about proxy quality. The more of it a correction removes, the worse the ranking because removal discards part of the signal but leaves the main sources of ranking noise, the noisy north star and the limited number of experiments, untouched. Archive-calibrated simulations, in which the correct ranking is known, confirm this even when every correction receives the true sampling covariance. Held-out real experiments, evaluated on disjoint customer halves so that shared error cannot bias the comparison, closely reproduce the predicted ordering (Spearman correlation 0.93). Correction can still pay off with more experiments, but the number needed rises steeply with the north star's noise. We map this crossover and give platform teams three inexpensive checks for deciding from their own archive whether and how strongly to correct.

Citation extraction

45
references
123
in-text mentions
45
distinct cited
0
self-citations
21,573
main-text words

appendix boundary found by appendix_command · 78% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Sigerson, L., Cunningham, T., Chou, W., Pandey, S., Stray, J., Yuan,… (2026) Evaluating for the long term: Learnings from industry0.69361100%
2Chou, W., Gray, C., Kallus, N., Bibaut, A., and Ejdemyr, S (2025) Evaluating decision rules across many weak experiments0.67211191%
3Tripuraneni, N., Richardson, L., D'Amour, A., Soriano, J., and Yadlo… (2024) Choosing a proxy metric from past experiments0.66910190%
4Bibaut, A., Chou, W., Ejdemyr, S., and Kallus, N (2024) Learning the covariance of treatment effects across many weak experiments0.66316188%
5Gazvoda, M. and Katsimerou, C (2024) Beyond correlation: A Bayesian multilevel model for effective proxy metric use in A/B tests0.6445180%
6Analytics at Meta (2022) Don't be seduced by the allure: A guide for how (not) to use proxy metrics in experiments0.64441100%
7Cunningham, T. and Kim, J (2022) Interpreting experiments with multiple outcomes0.64441100%
8Fleming, T. R. and DeMets, D. L (1996) Surrogate end points in clinical trials: Are we being misled?0.64441100%
9Bibaut, A., Kallus, N., Ejdemyr, S., and Zhao, M (2023) Long-term causal inference with imperfect surrogates using many weak experiments, proxies, and cross-fold moments0.58531100%
10Buyse, M., Molenberghs, G., Burzykowski, T., Renard, D., and Geys, H (2000) The validation of surrogate endpoints in meta-analyses of randomized experiments0.58531100%

Showing the top 10 of 45 scored citations.