Gözde Sert, Abhishek Chakrabortty, Anirban Bhattacharya
arXiv 22 Sep 2025 · Statistics — Methodology · publishedEconometrics and Statistics (2025)
arXiv:2509.17385 · PDF · DOI · OpenAlex · Extracted main text
Inference in semi-supervised (SS) settings has gained substantial attention in recent years due to increased relevance in modern big-data problems. In a typical SS setting, there is a much larger-sized unlabeled data, containing only observations of predictors, and a moderately sized labeled data containing observations for both an outcome and the set of predictors. Such data naturally arises when the outcome, unlike the predictors, is costly or difficult to obtain. One of the primary statistical objectives in SS settings is to explore whether parameter estimation can be improved by exploiting the unlabeled data. We propose a novel Bayesian method for estimating the population mean in SS settings. The approach yields estimators that are both efficient and optimal for estimation and inference. The method itself has several interesting artifacts. The central idea behind the method is to model certain summary statistics of the data in a targeted manner, rather than the entire raw data itself, along with a novel Bayesian notion of debiasing. Specifying appropriate summary statistics crucially relies on a debiased representation of the population mean that incorporates unlabeled data through a flexible nuisance function while also learning its estimation bias. Combined with careful usage of sample splitting, this debiasing approach mitigates the effect of bias due to slow rates or misspecification of the nuisance parameter from the posterior of the final parameter of interest, ensuring its robustness and efficiency. Concrete theoretical results, via Bernstein--von Mises theorems, are established, validating all claims, and are further supported through extensive numerical studies. To our knowledge, this is possibly the first work on Bayesian inference in SS settings, and its central ideas also apply more broadly to other Bayesian semi-parametric inference problems.
appendix boundary found by appendix_titled_section at “Supplementary Material” · 54% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/Debiased Machine Learning for Treatment and Structural Parameters | 0.822 | 12 | 2 | 83% |
| 2 | Aad W. van der Vaart (2000) Asymptotic Statistics, volume 3 | 0.781 | 7 | 2 | 71% |
| 3 | Valen E. Johnson and David Rossell (2012) Bayesian model selection in high-dimensional settings | 0.737 | 3 | 3 | 67% |
| 4 | Anru Zhang, Lawrence D. Brown, and T. Tony Cai (2019) Semi-supervised inference: General theory and estimation of means | 0.693 | 5 | 1 | 100% |
| 5 | Yuqian Zhang and Jelena Bradic (2022) High-dimensional semi-supervised learning: in search of optimal inference of the mean | 0.693 | 5 | 1 | 100% |
| 6 | Abhishek Chakrabortty and Tianxi Cai (2018) Efficient and adaptive linear regression in semi-supervised settings self | 0.644 | 4 | 1 | 100% |
| 7 | Abhishek Chakrabortty, Guorong Dai, and Raymond J Carroll (2022) Semi-supervised quantile estimation: Robust and efficient inference in high dimensional settings self | 0.644 | 3 | 2 | 67% |
| 8 | Alexandre B. Tsybakov (2009) Introduction to Nonparametric Estimation | 0.644 | 3 | 2 | 67% |
| 9 | Tong Zhang and Fernando J. Oles (2000) The value of unlabeled data for classification problems | 0.511 | 2 | 2 | 50% |
| 10 | Ashkan Ertefaie, Nima S. Hejazi, and Mark J. van der Laan (2022) Nonparametric inverse probability weighted estimators based on the highly adaptive lasso | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 63 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Bayesian semiparametric causal inference: Targeted doubly robust estimation of treatment effects | 0.481 | 18 | 4 |