arXiv 11 Jun 2026 · Statistics — Machine Learning
arXiv:2606.12892 · PDF · DOI · OpenAlex · Extracted main text
This study investigates semiparametric efficient estimation of causal and structural parameters in a semi-supervised setting. In our setting, unlabeled auxiliary regressors are available in addition to labeled observations consisting of outcomes and regressors. Our goal is to construct estimators of causal and structural parameters whose asymptotic variances are smaller than those of estimators constructed using only labeled data. We refer to this framework as prediction-powered causal inference (PPCI). We first derive the efficient influence function and the efficiency bound, which imply that the use of auxiliary regressors can attain a smaller asymptotic variance than the efficiency bound attainable from labeled observations alone. Then, by combining the efficient influence function with the debiased machine learning (DML) framework, we propose methods that we call DML-PPCI. If we construct an estimating-equation estimator, we refer to the method as EE-DML-PPCI; if we construct a targeted-learning estimator, we refer to the method as TMLE-DML-PPCI. The asymptotic variances of both estimators match our derived efficiency bound. In the construction of the estimators, estimation of the efficient influence function plays an important role. In our study, the efficient influence function is also a Neyman orthogonal score, which depends on the Riesz representer and the regression function. For Riesz representer estimation, we develop semi-supervised generalized Riesz regression with convergence rate guarantees.
appendix boundary found by appendix_command · 57% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Victor Chernozhukov, Whitney K. Newey, and Rahul Singh (2022) Automatic debiased machine learning of causal and structural effects | 1.000 | 7 | 3 | 100% |
| 2 | Masahiro Kato (2026) A unified framework for debiased machine learning: Riesz representer fitting under bregman divergence, 2026a self | 0.874 | 6 | 2 | 100% |
| 3 | Qingyuan Zhao (2019) Covariate balancing propensity score by tailored loss functions | 0.811 | 4 | 2 | 100% |
| 4 | Hidetoshi Shimodaira (2000) Improving predictive inference under covariate shift by weighting the log-likelihood function | 0.737 | 3 | 2 | 100% |
| 5 | Masatoshi Uehara, Masahiro Kato, and Shota Yasui (2020) Off-policy evaluation and learning for external validity under a covariate shift self | 0.644 | 4 | 1 | 100% |
| 6 | David Bruns-Smith, Oliver Dukes, Avi Feller, and Elizabeth L Ogburn (2025) Augmented balancing weights as linear regression | 0.644 | 2 | 2 | 100% |
| 7 | Masahiro Kato and Takeshi Teshima (2021) Non-negative bregman divergence minimization for deep direct density ratio estimation self | 0.644 | 2 | 2 | 100% |
| 8 | Larry Wasserman and John Lafferty (2007) Statistical analysis of semi-supervised regression | 0.644 | 2 | 2 | 100% |
| 9 | Ilker Demirel, Ahmed Alaa, Anthony Philippakis, and David Sontag (2024) Prediction-powered generalization of causal inferences | 0.511 | 2 | 1 | 100% |
| 10 | Masahiro Kato, Akihiro Oga, Wataru Komatsubara, and Ryo Inokuchi (2024) Active adaptive experimental design for treatment effect estimation with covariate choice self | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 51 scored citations.