EconBase
← All papers

When Do Surrogate Metrics Work? A Finite-Sample Comparison Under Realistic Failure Modes

Zihao Chen

arXiv 24 Sep 2026 · Statistics — Methodology

arXiv:2609.30528 · PDF · Extracted main text

Abstract

Online experiments must often be evaluated before long-term outcomes mature. Under rolling enrollment, these outcomes are observed only for early enrollees, while short-term surrogates are available for everyone. We compare seven estimators across eleven data-generating processes, spanning partial mediation, drift, outcome sparsity, and enrollment-time labeling, with up to $R = 2,000$ replications over more than 500 method-by-scenario cells. We find a sharp robustness-efficiency tradeoff: the surrogate index delivers large efficiency gains when surrogacy holds but its coverage collapses under violations, while PPI-family methods stay asymptotically valid under random labeling at smaller gains. We give the finite-sample variance of PPI++ in the all-units parameterization for a fixed predictor, a joint asymptotic distribution for the two estimators under cross-fitting, and a Hausman-type estimator-disagreement diagnostic, then quantify the detection-damage gap: in the partial-mediation design, where we locate both edges, a band of violations destroys surrogate-index coverage yet is too small to detect on most datasets. On the 64,000-customer Hillstrom experiment the diagnostic rarely flags a violation that biases the surrogate index, and a Cauchy-kernel hybrid of the two estimators inherits 10.7% relative bias; on the 14-million-user Criteo experiment the violation is detected, and subsampling traces detection turning on with scale as damage persists. PPI++ has limits: at a rare-conversion $n = 30,000$ Criteo subsample its empirical coverage is 85.0%. We recommend prespecifying PPI++ with the exact variance as the primary analysis under random labeling and adequate labeled outcome counts, reading the diagnostic as a warning, not a certificate, and treating the surrogate index as a sensitivity analysis.

Citation extraction

33
references
60
in-text mentions
33
distinct cited
0
self-citations
25,250
main-text words

appendix boundary found by appendix_command · 62% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Angelopoulos, Anastasios N. and Duchi, John C. and Zrnic, Tijana (2023) PPI++: Efficient Prediction-Powered Inference1.00063100%
2Ji, Wenlong and Lei, Lihua and Zrnic, Tijana (2025) Predictions as Surrogates: Revisiting Surrogate Outcomes in the Age of AI0.92843100%
3Leeb, Hannes and Pötscher, Benedikt M (2005) Model Selection and Inference: Facts and Fiction0.7374450%
4Angelopoulos, Anastasios N. and Bates, Stephen and Fannjiang, Clara… (2023) Prediction-Powered Inference0.73732100%
5Zrnic, Tijana and Candès, Emmanuel J (2024) Cross-Prediction-Powered Inference0.6443267%
6Athey, Susan and Chetty, Raj and Imbens, Guido W. and Kang, Hyunseung (2026) The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely0.64422100%
7Cassel, Claes-Magnus and Särndal, Carl-Erik and Wretman, Jan H (1976) Some Results on Generalized Difference Estimation and Generalized Regression Estimation for Finite Populations0.64422100%
8Guggenberger, Patrik (2010) The Impact of a Hausman Pretest on the Asymptotic Size of a Hypothesis Test0.64422100%
9Hausman, Jerry A (1978) Specification Tests in Econometrics0.64422100%
10Kilian, Valentin and Cortinovis, Stefano and Caron, Fran cois (2025) Anytime-Valid, Bayes-Assisted, Prediction-Powered Inference0.64422100%

Showing the top 10 of 33 scored citations.