EconBase
← All papers

Learning from Double Positive and Unlabeled Data for Potential-Customer Identification

Masahiro Kato, Yuki Ikeda, Kentaro Baba, Takashi Imai, Ryo Inokuchi

arXiv 31 May 2025 · Machine Learning

arXiv:2506.00436 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In this study, we propose a method for identifying potential customers in targeted marketing by applying learning from positive and unlabeled data (PU learning). We consider a scenario in which a company sells a product and can observe only the customers who purchased it. Decision-makers seek to market products effectively based on whether people have loyalty to the company. Individuals with loyalty are those who are likely to remain interested in the company even without additional advertising. Consequently, those loyal customers would likely purchase from the company if they are interested in the product. In contrast, people with lower loyalty may overlook the product or buy similar products from other companies unless they receive marketing attention. Therefore, by focusing marketing efforts on individuals who are interested in the product but do not have strong loyalty, we can achieve more efficient marketing. To achieve this goal, we consider how to learn, from limited data, a classifier that identifies potential customers who (i) have interest in the product and (ii) do not have loyalty to the company. Although our algorithm comprises a single-stage optimization, its objective function implicitly contains two losses derived from standard PU learning settings. For this reason, we refer to our approach as double PU learning. We verify the validity of the proposed algorithm through numerical experiments, confirming that it functions appropriately for the problem at hand.

Citation extraction

17
references
28
in-text mentions
17
distinct cited
3
self-citations
4,423
main-text words

appendix boundary found by appendix_command · 80% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Charles Elkan and Keith Noto (2008) Learning classifiers from only positive and unlabeled data0.8746467%
2Ryuichi Kiryo, Gang Niu, Marthinus Christoffel du Plessis, and Masas… (2017) Positive-unlabeled learning with non-negative risk estimator0.73732100%
3Marthinus Christoffel du Plessis, Gang. Niu, and Masashi Sugiyama (2015) Convex formulation for learning from positive and unlabeled data0.73732100%
4Gang Niu, Marthinus Christoffel du Plessis, Tomoya Sakai, Yao Ma, an… (2016) Theoretical comparisons of positive-unlabeled learning against positive-negative learning0.64422100%
5Tony Lancaster and Guido Imbens (1996) Case-control studies with contaminated controls0.51121100%
6Jessa Bekker and Jesse Davis (2018) Learning from positive and unlabeled data under the selected at random assumption0.40511100%
7Masahiro Kato, Liyuan Xu, Gang Niu, and Masashi Sugiyama (1809) Alternate estimation of a classifier and the class-prior from positive and unlabeled data, 2018 self0.40511100%
8Masahiro Kato and Takeshi Teshima (2021) Non-negative bregman divergence minimization for deep direct density ratio estimation self0.40511100%
9Jerzy Neyman (1923) Sur les applications de la theorie des probabilites aux experiences agricoles: Essai des principes0.40511100%
10A.W. van der Vaart (1998) Asymptotic Statistics0.40511100%

Showing the top 10 of 17 scored citations.