Sandro Provenzano
arXiv 6 Oct 2026 · Econometrics
arXiv:2610.08194 · PDF · Extracted main text
North-star metrics such as customer lifetime value are often too slow and noisy to decide a short A/B test. Teams therefore rely on a proxy metric, commonly chosen by how closely its effects tracked the north star's across past experiments. Validating that choice, or any method for making it, is hard: the only benchmark is the noisy north star, and the number of available past experiments is limited. In addition, proxy and north-star effects are estimated on the same customers, so their sampling errors are correlated. Recent work at major experimentation platforms removes this shared error as contamination, improving estimates of the true-effect covariance. Choosing a proxy, however, is a ranking problem, and a better estimate need not give a better ranking. We measure agreement free of shared error by estimating the two effects on disjoint random halves of each experiment's customers. In an archive of 262 experiments and 69 candidate proxies, the shared error ranks the candidates in a similar order to this agreement (Spearman correlation 0.65): it carries information about proxy quality. The more of it a correction removes, the worse the ranking because removal discards part of the signal but leaves the main sources of ranking noise, the noisy north star and the limited number of experiments, untouched. Archive-calibrated simulations, in which the correct ranking is known, confirm this even when every correction receives the true sampling covariance. Held-out real experiments, evaluated on disjoint customer halves so that shared error cannot bias the comparison, closely reproduce the predicted ordering (Spearman correlation 0.93). Correction can still pay off with more experiments, but the number needed rises steeply with the north star's noise. We map this crossover and give platform teams three inexpensive checks for deciding from their own archive whether and how strongly to correct.
appendix boundary found by appendix_command · 78% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Sigerson, L., Cunningham, T., Chou, W., Pandey, S., Stray, J., Yuan,… (2026) Evaluating for the long term: Learnings from industry | 0.693 | 6 | 1 | 100% |
| 2 | Chou, W., Gray, C., Kallus, N., Bibaut, A., and Ejdemyr, S (2025) Evaluating decision rules across many weak experiments | 0.672 | 11 | 1 | 91% |
| 3 | Tripuraneni, N., Richardson, L., D'Amour, A., Soriano, J., and Yadlo… (2024) Choosing a proxy metric from past experiments | 0.669 | 10 | 1 | 90% |
| 4 | Bibaut, A., Chou, W., Ejdemyr, S., and Kallus, N (2024) Learning the covariance of treatment effects across many weak experiments | 0.663 | 16 | 1 | 88% |
| 5 | Gazvoda, M. and Katsimerou, C (2024) Beyond correlation: A Bayesian multilevel model for effective proxy metric use in A/B tests | 0.644 | 5 | 1 | 80% |
| 6 | Analytics at Meta (2022) Don't be seduced by the allure: A guide for how (not) to use proxy metrics in experiments | 0.644 | 4 | 1 | 100% |
| 7 | Cunningham, T. and Kim, J (2022) Interpreting experiments with multiple outcomes | 0.644 | 4 | 1 | 100% |
| 8 | Fleming, T. R. and DeMets, D. L (1996) Surrogate end points in clinical trials: Are we being misled? | 0.644 | 4 | 1 | 100% |
| 9 | Bibaut, A., Kallus, N., Ejdemyr, S., and Zhao, M (2023) Long-term causal inference with imperfect surrogates using many weak experiments, proxies, and cross-fold moments | 0.585 | 3 | 1 | 100% |
| 10 | Buyse, M., Molenberghs, G., Burzykowski, T., Renard, D., and Geys, H (2000) The validation of surrogate endpoints in meta-analyses of randomized experiments | 0.585 | 3 | 1 | 100% |
Showing the top 10 of 45 scored citations.