Zihao Chen
arXiv 24 Sep 2026 · Statistics — Methodology
arXiv:2609.30528 · PDF · Extracted main text
Online experiments must often be evaluated before long-term outcomes mature. Under rolling enrollment, these outcomes are observed only for early enrollees, while short-term surrogates are available for everyone. We compare seven estimators across eleven data-generating processes, spanning partial mediation, drift, outcome sparsity, and enrollment-time labeling, with up to $R = 2,000$ replications over more than 500 method-by-scenario cells. We find a sharp robustness-efficiency tradeoff: the surrogate index delivers large efficiency gains when surrogacy holds but its coverage collapses under violations, while PPI-family methods stay asymptotically valid under random labeling at smaller gains. We give the finite-sample variance of PPI++ in the all-units parameterization for a fixed predictor, a joint asymptotic distribution for the two estimators under cross-fitting, and a Hausman-type estimator-disagreement diagnostic, then quantify the detection-damage gap: in the partial-mediation design, where we locate both edges, a band of violations destroys surrogate-index coverage yet is too small to detect on most datasets. On the 64,000-customer Hillstrom experiment the diagnostic rarely flags a violation that biases the surrogate index, and a Cauchy-kernel hybrid of the two estimators inherits 10.7% relative bias; on the 14-million-user Criteo experiment the violation is detected, and subsampling traces detection turning on with scale as damage persists. PPI++ has limits: at a rare-conversion $n = 30,000$ Criteo subsample its empirical coverage is 85.0%. We recommend prespecifying PPI++ with the exact variance as the primary analysis under random labeling and adequate labeled outcome counts, reading the diagnostic as a warning, not a certificate, and treating the surrogate index as a sensitivity analysis.
appendix boundary found by appendix_command · 62% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Angelopoulos, Anastasios N. and Duchi, John C. and Zrnic, Tijana (2023) PPI++: Efficient Prediction-Powered Inference | 1.000 | 6 | 3 | 100% |
| 2 | Ji, Wenlong and Lei, Lihua and Zrnic, Tijana (2025) Predictions as Surrogates: Revisiting Surrogate Outcomes in the Age of AI | 0.928 | 4 | 3 | 100% |
| 3 | Leeb, Hannes and Pötscher, Benedikt M (2005) Model Selection and Inference: Facts and Fiction | 0.737 | 4 | 4 | 50% |
| 4 | Angelopoulos, Anastasios N. and Bates, Stephen and Fannjiang, Clara… (2023) Prediction-Powered Inference | 0.737 | 3 | 2 | 100% |
| 5 | Zrnic, Tijana and Candès, Emmanuel J (2024) Cross-Prediction-Powered Inference | 0.644 | 3 | 2 | 67% |
| 6 | Athey, Susan and Chetty, Raj and Imbens, Guido W. and Kang, Hyunseung (2026) The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely | 0.644 | 2 | 2 | 100% |
| 7 | Cassel, Claes-Magnus and Särndal, Carl-Erik and Wretman, Jan H (1976) Some Results on Generalized Difference Estimation and Generalized Regression Estimation for Finite Populations | 0.644 | 2 | 2 | 100% |
| 8 | Guggenberger, Patrik (2010) The Impact of a Hausman Pretest on the Asymptotic Size of a Hypothesis Test | 0.644 | 2 | 2 | 100% |
| 9 | Hausman, Jerry A (1978) Specification Tests in Econometrics | 0.644 | 2 | 2 | 100% |
| 10 | Kilian, Valentin and Cortinovis, Stefano and Caron, Fran cois (2025) Anytime-Valid, Bayes-Assisted, Prediction-Powered Inference | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 33 scored citations.