paper-with-me

홈 › Papers

On the Rates of Convergence from Surrogate Risk Minimizers to the Bayes Optimal Classifier

2018-02-11 · Jingwei Zhang, Tongliang Liu, DaCheng Tao

We study the rates of convergence from empirical surrogate risk minimizers to the Bayes optimal classifier. Specifically, we introduce the notion of \emph{consistency intensity} to characterize a surrogate loss function and exploit this notion to obtain the rate of convergence from an empirical surrogate risk minimizer to the Bayes optimal classifier, enabling fair comparisons of the excess risks of different surrogate risk minimizers. The main result of the paper has practical implications including (1) showing that hinge loss is superior to logistic and exponential loss in the sense that its empirical minimizer converges faster to the Bayes optimal classifier and (2) guiding to modify surrogate loss functions to accelerate the convergence to the Bayes optimal classifier.

📄 PDF Abstract BibTeX arXiv:1802.03688

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adversarial Consistency and the Uniqueness of the Adversarial Bayes Classifier

2024-04-26 · Natalie S. Frank

Minimizing an adversarial surrogate risk is a common technique for learning robust classifiers. Prior work showed that convex surrogate losses are not statistically consistent in the adversarial context -- or in other wo…

Classification

Proper losses regret at least 1/2-order

2024-07-15 · Han Bao, Asuka Takatsu

A fundamental challenge in machine learning is the choice of a loss as it characterizes our learning task, is minimized in the training phase, and serves as an evaluation criterion for estimators. Proper losses are commo…

Establishing Linear Surrogate Regret Bounds for Convex Smooth Losses via Convolutional Fenchel-Young Losses

2025-05-14 · Yuzhou Cao, Han Bao, Lei Feng, Bo An

Surrogate regret bounds, also known as excess risk bounds, bridge the gap between the convergence rates of surrogate and target losses, with linear bounds favorable for their lossless regret transfer. While convex smooth…

Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks

2024-02-07 · Arnulf Jentzen, Adrian Riekert

Stochastic gradient descent (SGD) optimization methods such as the plain vanilla SGD method and the popular Adam optimizer are nowadays the method of choice in the training of artificial neural networks (ANNs). Despite t…

Exponential Convergence Rates of Classification Errors on Learning with SGD and Random Features

2019-11-13 · Shingo Yashima, Atsushi Nitanda, Taiji Suzuki

Although kernel methods are widely used in many learning problems, they have poor scalability to large datasets. To address this problem, sketching and stochastic gradient methods are the most commonly used techniques to…

Binary ClassificationClassificationGeneral Classification