arXiv 14 Nov 2025 · Mathematics — Statistics Theory
arXiv:2511.11003 · PDF · DOI · OpenAlex · Extracted main text
Distribution shift between the training domain and the test domain poses a key challenge for modern machine learning. An extensively studied instance is the covariate shift, where the marginal distribution of covariates differs across domains, while the conditional distribution of outcome remains the same. The doubly-robust (DR) estimator, recently introduced by \cite{kato2023double}, combines the density ratio estimation with a pilot regression model and demonstrates asymptotic normality and $\sqrt{n}$-consistency, even when the pilot estimates converge slowly. However, the prior arts has focused exclusively on deriving asymptotic results and has left open the question of non-asymptotic guarantees for the DR estimator. This paper establishes the first non-asymptotic learning bounds for the DR covariate shift adaptation. Our main contributions are two-fold: (\romannumeral 1) We establish structure-agnostic high-probability upper bounds on the excess target risk of the DR estimator that depend only on the $L^2$-errors of the pilot estimates and the Rademacher complexity of the model class, without assuming specific procedures to obtain the pilot estimate, and (\romannumeral 2) under well-specified parameterized models, we analyze the DR covariate shift adaptation based on modern techniques for non-asymptotic analysis of MLE, whose key terms governed by the Fisher information mismatch term between the source and target distributions. Together, these findings bridge asymptotic efficiency properties and a finite-sample out-of-distribution generalization bounds, providing a comprehensive theoretical underpinnings for the DR covariate shift adaptation.
appendix boundary found by appendix_command · 27% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Masahiro Kato, Kota Matsui, and Ryo Inokuchi (2023) Double debiased covariate shift adaptation robust to density-ratio estimation | 1.000 | 5 | 4 | 100% |
| 2 | Hidetoshi Shimodaira (2000) Improving predictive inference under covariate shift by weighting the log-likelihood function | 0.928 | 4 | 3 | 100% |
| 3 | James M Robins and Andrea Rotnitzky (1995) Semiparametric efficiency in multivariate regression models with missing data | 0.843 | 3 | 3 | 100% |
| 4 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2017) Double/debiased/neyman machine learning of treatment effects | 0.737 | 3 | 2 | 100% |
| 5 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters | 0.737 | 3 | 2 | 100% |
| 6 | Cong Ma, Reese Pathak, and Martin J Wainwright (2023) Optimally tackling covariate shift in rkhs-based nonparametric regression self | 0.737 | 3 | 2 | 100% |
| 7 | Sivaraman Balakrishnan, Edward H Kennedy, and Larry Wasserman (2023) The fundamental limits of structure-agnostic functional estimation | 0.644 | 2 | 2 | 100% |
| 8 | Matteo Bonvini, Edward H Kennedy, Oliver Dukes, and Sivaraman Balakr… (2024) Doubly-robust inference and optimality in structure-agnostic models with smoothness | 0.644 | 2 | 2 | 100% |
| 9 | Victor Chernozhukov, Juan Carlos Escanciano, Hidehiko Ichimura, Whit… (2022) Locally robust semiparametric estimation | 0.644 | 2 | 2 | 100% |
| 10 | Victor Chernozhukov, Whitney K Newey, and Rahul Singh (2023) A simple and general debiased machine learning theorem with finite-sample guarantees | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 87 scored citations.