EconBase
← All papers

Learning bounds for doubly-robust covariate shift adaptation

Jeonghwan Lee, Cong Ma

arXiv 14 Nov 2025 · Mathematics — Statistics Theory

arXiv:2511.11003 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Distribution shift between the training domain and the test domain poses a key challenge for modern machine learning. An extensively studied instance is the covariate shift, where the marginal distribution of covariates differs across domains, while the conditional distribution of outcome remains the same. The doubly-robust (DR) estimator, recently introduced by \cite{kato2023double}, combines the density ratio estimation with a pilot regression model and demonstrates asymptotic normality and $\sqrt{n}$-consistency, even when the pilot estimates converge slowly. However, the prior arts has focused exclusively on deriving asymptotic results and has left open the question of non-asymptotic guarantees for the DR estimator. This paper establishes the first non-asymptotic learning bounds for the DR covariate shift adaptation. Our main contributions are two-fold: (\romannumeral 1) We establish structure-agnostic high-probability upper bounds on the excess target risk of the DR estimator that depend only on the $L^2$-errors of the pilot estimates and the Rademacher complexity of the model class, without assuming specific procedures to obtain the pilot estimate, and (\romannumeral 2) under well-specified parameterized models, we analyze the DR covariate shift adaptation based on modern techniques for non-asymptotic analysis of MLE, whose key terms governed by the Fisher information mismatch term between the source and target distributions. Together, these findings bridge asymptotic efficiency properties and a finite-sample out-of-distribution generalization bounds, providing a comprehensive theoretical underpinnings for the DR covariate shift adaptation.

Citation extraction

87
references
124
in-text mentions
87
distinct cited
4
self-citations
10,756
main-text words

appendix boundary found by appendix_command · 27% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Masahiro Kato, Kota Matsui, and Ryo Inokuchi (2023) Double debiased covariate shift adaptation robust to density-ratio estimation1.00054100%
2Hidetoshi Shimodaira (2000) Improving predictive inference under covariate shift by weighting the log-likelihood function0.92843100%
3James M Robins and Andrea Rotnitzky (1995) Semiparametric efficiency in multivariate regression models with missing data0.84333100%
4Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2017) Double/debiased/neyman machine learning of treatment effects0.73732100%
5Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters0.73732100%
6Cong Ma, Reese Pathak, and Martin J Wainwright (2023) Optimally tackling covariate shift in rkhs-based nonparametric regression self0.73732100%
7Sivaraman Balakrishnan, Edward H Kennedy, and Larry Wasserman (2023) The fundamental limits of structure-agnostic functional estimation0.64422100%
8Matteo Bonvini, Edward H Kennedy, Oliver Dukes, and Sivaraman Balakr… (2024) Doubly-robust inference and optimality in structure-agnostic models with smoothness0.64422100%
9Victor Chernozhukov, Juan Carlos Escanciano, Hidehiko Ichimura, Whit… (2022) Locally robust semiparametric estimation0.64422100%
10Victor Chernozhukov, Whitney K Newey, and Rahul Singh (2023) A simple and general debiased machine learning theorem with finite-sample guarantees0.64422100%

Showing the top 10 of 87 scored citations.