Masahiro Kato, Kota Matsui, Ryo Inokuchi
arXiv 25 Oct 2023 · Statistics — Methodology
arXiv:2310.16638 · PDF · DOI · OpenAlex · Extracted main text
Consider a scenario where we have access to train data with both covariates and outcomes while test data only contains covariates. In this scenario, our primary aim is to predict the missing outcomes of the test data. With this objective in mind, we train parametric regression models under a covariate shift, where covariate distributions are different between the train and test data. For this problem, existing studies have proposed covariate shift adaptation via importance weighting using the density ratio. This approach averages the train data losses, each weighted by an estimated ratio of the covariate densities between the train and test data, to approximate the test-data risk. Although it allows us to obtain a test-data risk minimizer, its performance heavily relies on the accuracy of the density ratio estimation. Moreover, even if the density ratio can be consistently estimated, the estimation errors of the density ratio also yield bias in the estimators of the regression model's parameters of interest. To mitigate these challenges, we introduce a doubly robust estimator for covariate shift adaptation via importance weighting, which incorporates an additional estimator for the regression function. Leveraging double machine learning techniques, our estimator reduces the bias arising from the density ratio estimation errors. We demonstrate the asymptotic distribution of the regression parameter estimator. Notably, our estimator remains consistent if either the density ratio estimator or the regression function is consistent, showcasing its robustness against potential errors in density ratio estimation. Finally, we confirm the soundness of our proposed method via simulation studies.
appendix boundary found by appendix_command · 55% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Shimodaira (2000) `Improving predictive inference under covariate shift by weighting the log-likelihood function', Journal of Statistical Planning… | 0.969 | 11 | 4 | 91% |
| 2 | Kato \ Teshima (2020) `Non-negative bregman divergence minimization for deep direct density ratio estimation' | 0.928 | 4 | 3 | 100% |
| 3 | Kanamori, Hido \ Sugiyama (2009) `A least-squares approach to direct importance estimation', Journal of Machine Learning Research 10(Jul.), 1391–1445 | 0.843 | 3 | 3 | 100% |
| 4 | Kanamori, Suzuki \ Sugiyama (2012) `Statistical analysis of kernel-based least-squares density-ratio estimation', Machine Learning 86(3), 335–367 | 0.737 | 3 | 2 | 100% |
| 5 | Sugiyama, Nakajima, Kashima, Buenau \ Kawanabe (2008) Direct importance estimation with model selection and its application to covariate shift adaptation, in `NeuIPS', pp. 1433–1440 | 0.737 | 3 | 2 | 100% |
| 6 | Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, Newey \ Robins (2018) `Double/debiased machine learning for treatment and structural parameters', Econometrics Journal 21, C1–C68 | 0.644 | 2 | 2 | 100% |
| 7 | Reddi, Póczos \ Smola (2015) Doubly robust covariate shift correction, in `AAAI', AAAI Press, pp. 2949–2955 | 0.644 | 2 | 2 | 100% |
| 8 | Sugiyama, Suzuki \ Kanamori (2012) Density Ratio Estimation in Machine Learning, 1st edn, Cambridge University Press, New York, NY, USA | 0.644 | 2 | 2 | 100% |
| 9 | Uehara, Kato \ Yasui (2020) Off-policy evaluation and learning for external validity under a covariate shift, in `Advances in Neural Information Processing… | 0.511 | 5 | 2 | 20% |
| 10 | van der Vaart (1998) Asymptotic statistics, Cambridge University Press, Cambridge, UK | 0.511 | 2 | 2 | 50% |
Showing the top 10 of 53 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Learning bounds for doubly-robust covariate shift adaptation | 1.000 | 5 | 4 |