Ruicheng Ao, Hongyu Chen, Haoyang Liu, David Simchi-Levi, Will Wei Sun
arXiv 29 Jan 2026 · Machine Learning
arXiv:2601.21470 · PDF · DOI · OpenAlex · Extracted main text
We study semi-supervised stochastic optimization when labeled data is scarce but predictions from pre-trained models are available. PPI and SVRG both reduce variance through control variates -- PPI uses predictions, SVRG uses reference gradients. We show they are mathematically equivalent and develop PPI-SVRG, which combines both. Our convergence bound decomposes into the standard SVRG rate plus an error floor from prediction uncertainty. The rate depends only on loss geometry; predictions affect only the neighborhood size. When predictions are perfect, we recover SVRG exactly. When predictions degrade, convergence remains stable but reaches a larger neighborhood. Experiments confirm the theory: PPI-SVRG reduces MSE by 43--52% under label scarcity on mean estimation benchmarks and improves test accuracy by 2.7--2.9 percentage points on MNIST with only 10% labeled data.
appendix boundary found by appendix_command · 49% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Rie Johnson and Tong Zhang (2013) Accelerating stochastic gradient descent using predictive variance reduction | 1.000 | 5 | 4 | 100% |
| 2 | Anastasios N Angelopoulos, Stephen Bates, Clara Fannjiang, Michael I… (2023) Prediction-powered inference | 0.928 | 5 | 5 | 80% |
| 3 | Zeyuan Allen-Zhu and Yang Yuan (2016) Improved svrg for non-strongly-convex or sum-of-non-convex objectives | 0.693 | 6 | 3 | 33% |
| 4 | Frank Schneider, Lukas Balles, and Philipp Hennig (2019) Deepobs: A deep learning optimizer benchmark suite | 0.511 | 4 | 2 | 25% |
| 5 | Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika… (2018) Catboost: unbiased boosting with categorical features | 0.511 | 2 | 2 | 50% |
| 6 | Anastasios N Angelopoulos, John C Duchi, and Tijana Zrnic (2023) Ppi++: Efficient prediction-powered inference | 0.405 | 1 | 1 | 100% |
| 7 | Ruicheng Ao, Hongyu Chen, and David Simchi-Levi (2024) Prediction-guided active experiments self | 0.405 | 1 | 1 | 100% |
| 8 | Aaron Defazio, Francis Bach, and Simon Lacoste-Julien (2014) Saga: A fast incremental gradient method with support for non-strongly convex composite objectives | 0.405 | 1 | 1 | 100% |
| 9 | Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang (2018) Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator | 0.405 | 1 | 1 | 100% |
| 10 | Neal Jean, Marshall Burke, Michael Xie, W Matthew Davis, David B Lob… (2016) Combining satellite imagery and machine learning to predict poverty | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 21 scored citations.