EconBase
← All papers

Policy-Oriented Binary Classification: Improving (KD-)CART Final Splits for Subpopulation Targeting

Lei Bill Wang, Zhenbang Jiao, Fangyi Wang

arXiv 20 Feb 2025 · Statistics — Machine Learning

arXiv:2502.15072 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Policymakers often use recursive binary split rules to partition populations based on binary outcomes and target subpopulations whose probability of the binary event exceeds a threshold. We call such problems Latent Probability Classification (LPC). Practitioners typically employ Classification and Regression Trees (CART) for LPC. We prove that in the context of LPC, classic CART and the knowledge distillation method, whose student model is a CART (referred to as KD-CART), are suboptimal. We propose Maximizing Distance Final Split (MDFS), which generates split rules that strictly dominate CART/KD-CART under the unique intersect assumption. MDFS identifies the unique best split rule, is consistent, and targets more vulnerable subpopulations than CART/KD-CART. To relax the unique intersect assumption, we additionally propose Penalized Final Split (PFS) and weighted Empirical risk Final Split (wEFS). Through extensive simulation studies, we demonstrate that the proposed methods predominantly outperform CART/KD-CART. When applied to real-world datasets, MDFS generates policies that target more vulnerable subpopulations than the CART/KD-CART.

Citation extraction

50
references
56
in-text mentions
50
distinct cited
1
self-citations
6,658
main-text words

appendix boundary found by appendix_command · 34% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Oluwasanmi O Koyejo, Nagarajan Natarajan, Pradeep K Ravikumar, and I… (2014) Consistent binary classification with generalized performance metrics0.73732100%
2Susan Athey and Guido Imbens (2016) Recursive partitioning for heterogeneous causal effects0.5112250%
3Ye Nan, Kian Ming Chai, Wee Sun Lee, and Hai Leong Chieu (2012) Optimizing f-measure: A tale of two approaches0.51121100%
4Monica Andini, Emanuele Ciani, Guido de Blasio, Alessio D'Ignazio, a… (2018) Targeting with machine learning: An application to a tax rebate program in italy0.40511100%
5Susan Athey and Stefan Wager (2021) Policy learning with observational data0.40511100%
6Martin Atzmueller (2015) Subgroup discovery0.40511100%
7Andrii Babii, Xi Chen, Eric Ghysels, and Rohit Kumar (2024) Binary choice with asymmetric loss in a data-rich environment: Theory and an application to racial justice0.40511100%
8Guy Blanc, Jane Lange, and Li-Yang Tan (2020) Provable guarantees for decision tree induction: the agnostic setting0.40511100%
9Leo Breiman (1996) Bagging predictors0.40511100%
10Tri Dao, Govinda M Kamath, Vasilis Syrgkanis, and Lester Mackey (2021) Knowledge distillation as semiparametric inference0.40511100%

Showing the top 10 of 50 scored citations.