EconBase
← All papers

Constrained Classification and Policy Learning

Toru Kitagawa, Shosei Sakaguchi, Aleksey Tetenov

arXiv 24 Jun 2021 · Econometrics · 27 citations (OpenAlex)

arXiv:2106.12886 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Modern machine learning approaches to classification, including AdaBoost, support vector machines, and deep neural networks, utilize surrogate loss techniques to circumvent the computational complexity of minimizing empirical classification risk. These techniques are also useful for causal policy learning problems, since estimation of individualized treatment rules can be cast as a weighted (cost-sensitive) classification problem. Consistency of the surrogate loss approaches studied in Zhang (2004) and Bartlett et al. (2006) crucially relies on the assumption of correct specification, meaning that the specified set of classifiers is rich enough to contain a first-best classifier. This assumption is, however, less credible when the set of classifiers is constrained by interpretability or fairness, leaving the applicability of surrogate loss based algorithms unknown in such second-best scenarios. This paper studies consistency of surrogate loss procedures under a constrained set of classifiers without assuming correct specification. We show that in the setting where the constraint restricts the classifier's prediction set only, hinge losses (i.e., $\ell_1$-support vector machines) are the only surrogate losses that preserve consistency in second-best scenarios. If the constraint additionally restricts the functional form of the classifier, consistency of a surrogate loss approach is not guaranteed even with hinge loss. We therefore characterize conditions for the constrained set of classifiers that can guarantee consistency of hinge risk minimizing classifiers. Exploiting our theoretical results, we develop robust and computationally attractive hinge loss based procedures for a monotone classification problem.

Citation extraction

66
references
141
in-text mentions
66
distinct cited
6
self-citations
20,567
main-text words

appendix boundary found by appendix_command · 49% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Zhang, T (2004) Statistical behavior and consistency of classification methods based on convex risk minimization1.000126100%
2Nguyen, X., M. J. Wainwright, and M. I. Jordan (2009) On surrogate loss functions and f-divergences1.000113100%
3Bartlett, P. L., M. I. Jordan, and J. D. McAuliffe (2006) Convexity, classification, and risk bounds0.95315487%
4Mbakop, E. and M. Tabord-Meehan (2021) Model selection for treatment choice: Penalized welfare maximization0.86011464%
5Dudley, R. M (1999) Uniform Central Limit Theorems0.8434375%
6Dwork, C., M. Hardt, T. Pitassi, O. Reingold, and R. Zemel (2012) Fairness through awareness, in0.73732100%
7Karlan, D., S. Mullainathan, and B. N. Roth (2019) Debt traps? Market vendors and moneylender debt in India and the Philippines0.69361100%
8Athey, S. and S. Wager (2021) Policy learning with observational data0.64441100%
9Kitagawa, T. and A. Tetenov (2018) Who should be treated? Empirical welfare maximization methods for treatment choice self0.64441100%
10Chen, C. C. and S. T. Li (2014) Credit rating with a monotonicity-constrained support vector machine model0.64422100%

Showing the top 10 of 66 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Nonparametric Uniform Inference in Binary Classification and Policy Values0.73732
2Estimation of Optimal Dynamic Treatment Assignment Rules under Policy Constraints0.51121
3Who Should Get Vaccinated? Individualized Allocation of Vaccines Over SIR Network0.40511
4Evidence Aggregation for Treatment Choice0.40511
5Orthogonal Policy Learning Under Ambiguity0.40511
6Stochastic treatment choice with empirical welfare updating0.40511
7Individualized Treatment Allocation in Sequential Network Games0.40511
8Econometrics of Machine Learning Methods in Economic Forecasting0.40511
9Policy Learning with Distributional Welfare0.40511
10Decision Theory for Treatment Choice Problems with Partial Identification0.40511