paper-with-me

홈 › Papers

Relaxed Clipping: A Global Training Method for Robust Regression and Classification

2010-12-01 · NeurIPS 2010 12 · Min Yang, Linli Xu, Martha White, Dale Schuurmans, Yao-Liang Yu

Robust regression and classification are often thought to require non-convex loss functions that prevent scalable, global training. However, such a view neglects the possibility of reformulated training methods that can yield practically solvable alternatives. A natural way to make a loss function more robust to outliers is to truncate loss values that exceed a maximum threshold. We demonstrate that a relaxation of this form of ``loss clipping'' can be made globally solvable and applicable to any standard loss while guaranteeing robustness against outliers. We present a generic procedure that can be applied to standard loss functions and demonstrate improved robustness in regression and classification problems.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral Classificationregression

Similar Papers 제목 키워드 기반

Variance-reduced Clipping for Non-convex Optimization

2023-03-02 · Amirhossein Reisizadeh, Haochuan Li, Subhro Das, Ali Jadbabaie

Gradient clipping is a standard training technique used in deep learning applications such as large-scale language modeling to mitigate exploding gradients. Recent experimental studies have demonstrated a fairly special …

Language ModelingLanguage Modelling

EPISODE: Episodic Gradient Clipping with Periodic Resampled Corrections for Federated Learning with Heterogeneous Data

2023-02-14 · Michael Crawshaw, Yajie Bao, Mingrui Liu

Gradient clipping is an important technique for deep neural networks with exploding gradients, such as recurrent neural networks. Recent studies have shown that the loss functions of these networks do not satisfy the con…

Federated Learning

Robustness to Unbounded Smoothness of Generalized SignSGD

2022-08-23 · Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei zhang 외

Traditional analyses in non-convex optimization typically rely on the smoothness assumption, namely requiring the gradients to be Lipschitz. However, recent evidence shows that this smoothness condition does not capture …

A Communication-Efficient Distributed Gradient Clipping Algorithm for Training Deep Neural Networks

2022-05-10 · Mingrui Liu, Zhenxun Zhuang, Yunwei Lei, Chunyang Liao

In distributed training of deep neural networks, people usually run Stochastic Gradient Descent (SGD) or its variants on each machine and communicate with other machines periodically. However, SGD might converge slowly i…

Federated Learning

Fisher Discriminative Least Squares Regression for Image Classification

2019-03-19 · Zhe Chen, Xiao-Jun Wu, Josef Kittler

Discriminative least squares regression (DLSR) has been shown to achieve promising performance in multi-class image classification tasks. Its key idea is to force the regression labels of different classes to move in opp…

ClassificationFace RecognitionGeneral Classificationimage-classification+2