Ruicheng Ao, Hongyu Chen, David Simchi-Levi
arXiv 18 Nov 2024 · Statistics — Machine Learning · 2 citations (OpenAlex)
arXiv:2411.12036 · PDF · DOI · OpenAlex · Extracted main text
In this work, we introduce a new framework for active experimentation, the Prediction-Guided Active Experiment (PGAE), which leverages predictions from an existing machine learning model to guide sampling and experimentation. Specifically, at each time step, an experimental unit is sampled according to a designated sampling distribution, and the actual outcome is observed based on an experimental probability. Otherwise, only a prediction for the outcome is available. We begin by analyzing the non-adaptive case, where full information on the joint distribution of the predictor and the actual outcome is assumed. For this scenario, we derive an optimal experimentation strategy by minimizing the semi-parametric efficiency bound for the class of regular estimators. We then introduce an estimator that meets this efficiency bound, achieving asymptotic optimality. Next, we move to the adaptive case, where the predictor is continuously updated with newly sampled data. We show that the adaptive version of the estimator remains efficient and attains the same semi-parametric bound under certain regularity assumptions. Finally, we validate PGAE's performance through simulations and a semi-synthetic experiment using data from the US Census Bureau. The results underscore the PGAE framework's effectiveness and superiority compared to other existing methods.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Angelopoulos AN, Bates S, Fannjiang C, Jordan MI, Zrnic T (2023) a) Prediction-powered inference | 0.811 | 4 | 2 | 100% |
| 2 | Van der Vaart AW (2000) Asymptotic statistics | 0.811 | 4 | 2 | 100% |
| 3 | Kato M, Oga A, Komatsubara W, Inokuchi R (2024) Active adaptive experimental design for treatment effect estimation with covariate choice | 0.737 | 3 | 2 | 100% |
| 4 | Kennedy EH (2022) Semiparametric doubly robust targeted double machine learning: a review | 0.737 | 3 | 2 | 100% |
| 5 | Zrnic T, Candès EJ (2024) a) Active statistical inference | 0.737 | 3 | 2 | 100% |
| 6 | Chernozhukov V, Chetverikov D, Demirer M, Duflo E, Hansen C, Newey W… (2018) Double/debiased machine learning for treatment and structural parameters | 0.644 | 2 | 2 | 100% |
| 7 | Hamilton JD (2020) Time series analysis | 0.511 | 2 | 1 | 100% |
| 8 | Li H, Zhao G, Johari R, Weintraub GY (2022) Interference, bias, and variance in two-sided marketplace experimentation: Guidance for platforms | 0.511 | 2 | 1 | 100% |
| 9 | Angelopoulos AN, Duchi JC, Zrnic T (2023) b) Ppi++: Efficient prediction-powered inference | 0.405 | 1 | 1 | 100% |
| 10 | Ao R, Chen H, Simchi-Levi D, Zhu F (2024) Online local false discovery rate control: A resource allocation approach | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 48 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | PPI-SVRG: Unifying Prediction-Powered Inference and Variance Reduction for Semi-Supervised Optimization | 0.405 | 1 | 1 |