EconBase
← All papers

PUATE: Efficient Average Treatment Effect Estimation from Treated (Positive) and Unlabeled Units

Masahiro Kato, Fumiaki Kozai, Ryo Inokuchi

arXiv 31 Jan 2025 · Machine Learning

arXiv:2501.19345 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The estimation of average treatment effects (ATEs), defined as the difference in expected outcomes between treatment and control groups, is a central topic in causal inference. This study develops semiparametric efficient estimators for ATE in a setting where only a treatment group and an unlabeled group, consisting of units whose treatment status is unknown, are observed. This scenario constitutes a variant of learning from positive and unlabeled data (PU learning) and can be viewed as a special case of ATE estimation with missing data. For this setting, we derive the semiparametric efficiency bounds, which characterize the lowest achievable asymptotic variance for regular estimators. We then construct semiparametric efficient ATE estimators that attain these bounds. Our results contribute to the literature on causal inference with missing data and weakly supervised learning.

Citation extraction

76
references
147
in-text mentions
76
distinct cited
7
self-citations
10,460
main-text words

appendix boundary found by appendix_command · 40% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Aad W. van der Vaart (1998) Asymptotic Statistics0.8434375%
2Charles Elkan and Keith Noto (2008) Learning classifiers from only positive and unlabeled data0.70717935%
3Marthinus Christoffel du Plessis, Gang. Niu, and Masashi Sugiyama (2015) Convex formulation for learning from positive and unlabeled data0.6939733%
4Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters0.6444250%
5Gang Niu, Marthinus Christoffel du Plessis, Tomoya Sakai, Yao Ma, an… (2016) Theoretical comparisons of positive-unlabeled learning against positive-negative learning0.64422100%
6Donald B. Rubin (1974) Estimating causal effects of treatments in randomized and nonrandomized studies0.64422100%
7Guido W. Imbens and Tony Lancaster (1996) Efficient estimation and stratified sampling0.5853333%
8Jeffrey M. Wooldridge (2001) Asymptotic properties of weighted m-estimation for standard stratified samples0.5853333%
9Heejung Bang and James M. Robins (2005) Doubly robust estimation in missing data and causal inference models0.5113233%
10Jessa Bekker and Jesse Davis (2018) Learning from positive and unlabeled data under the selected at random assumption0.5113233%

Showing the top 10 of 76 scored citations.