Toru Kitagawa, Shosei Sakaguchi, Aleksey Tetenov
arXiv 24 Jun 2021 · Econometrics · 27 citations (OpenAlex)
arXiv:2106.12886 · PDF · DOI · OpenAlex · Extracted main text
Modern machine learning approaches to classification, including AdaBoost, support vector machines, and deep neural networks, utilize surrogate loss techniques to circumvent the computational complexity of minimizing empirical classification risk. These techniques are also useful for causal policy learning problems, since estimation of individualized treatment rules can be cast as a weighted (cost-sensitive) classification problem. Consistency of the surrogate loss approaches studied in Zhang (2004) and Bartlett et al. (2006) crucially relies on the assumption of correct specification, meaning that the specified set of classifiers is rich enough to contain a first-best classifier. This assumption is, however, less credible when the set of classifiers is constrained by interpretability or fairness, leaving the applicability of surrogate loss based algorithms unknown in such second-best scenarios. This paper studies consistency of surrogate loss procedures under a constrained set of classifiers without assuming correct specification. We show that in the setting where the constraint restricts the classifier's prediction set only, hinge losses (i.e., $\ell_1$-support vector machines) are the only surrogate losses that preserve consistency in second-best scenarios. If the constraint additionally restricts the functional form of the classifier, consistency of a surrogate loss approach is not guaranteed even with hinge loss. We therefore characterize conditions for the constrained set of classifiers that can guarantee consistency of hinge risk minimizing classifiers. Exploiting our theoretical results, we develop robust and computationally attractive hinge loss based procedures for a monotone classification problem.
appendix boundary found by appendix_command · 49% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Zhang, T (2004) Statistical behavior and consistency of classification methods based on convex risk minimization | 1.000 | 12 | 6 | 100% |
| 2 | Nguyen, X., M. J. Wainwright, and M. I. Jordan (2009) On surrogate loss functions and f-divergences | 1.000 | 11 | 3 | 100% |
| 3 | Bartlett, P. L., M. I. Jordan, and J. D. McAuliffe (2006) Convexity, classification, and risk bounds | 0.953 | 15 | 4 | 87% |
| 4 | Mbakop, E. and M. Tabord-Meehan (2021) Model selection for treatment choice: Penalized welfare maximization | 0.860 | 11 | 4 | 64% |
| 5 | Dudley, R. M (1999) Uniform Central Limit Theorems | 0.843 | 4 | 3 | 75% |
| 6 | Dwork, C., M. Hardt, T. Pitassi, O. Reingold, and R. Zemel (2012) Fairness through awareness, in | 0.737 | 3 | 2 | 100% |
| 7 | Karlan, D., S. Mullainathan, and B. N. Roth (2019) Debt traps? Market vendors and moneylender debt in India and the Philippines | 0.693 | 6 | 1 | 100% |
| 8 | Athey, S. and S. Wager (2021) Policy learning with observational data | 0.644 | 4 | 1 | 100% |
| 9 | Kitagawa, T. and A. Tetenov (2018) Who should be treated? Empirical welfare maximization methods for treatment choice self | 0.644 | 4 | 1 | 100% |
| 10 | Chen, C. C. and S. T. Li (2014) Credit rating with a monotonicity-constrained support vector machine model | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 66 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.