paper-with-me

홈 › Papers

F10-SGD: Fast Training of Elastic-net Linear Models for Text Classification and Named-entity Recognition

2019-02-27 · Stanislav Peshterliev, Alexander Hsieh, Imre Kiss

Voice-assistants text classification and named-entity recognition (NER) models are trained on millions of example utterances. Because of the large datasets, long training time is one of the bottlenecks for releasing improved models. In this work, we develop F10-SGD, a fast optimizer for text classification and NER elastic-net linear models. On internal datasets, F10-SGD provides 4x reduction in training time compared to the OWL-QN optimizer without loss of accuracy or increase in model size. Furthermore, we incorporate biased sampling that prioritizes harder examples towards the end of the training. As a result, in addition to faster training, we were able to obtain statistically significant accuracy improvements for NER. On public datasets, F10-SGD obtains 22% faster training time compared to FastText for text classification. And, 4x reduction in training time compared to CRFSuite OWL-QN for NER.

📄 PDF Abstract BibTeX arXiv:1902.10649

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral Classificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERtext-classificationText Classification

Similar Papers 제목 키워드 기반

A generalized linear joint trained framework for semi-supervised learning of sparse features

2020-06-02 · Juan C. Laria, Line H. Clemmensen, Bjarne K. Ersbøll

The elastic-net is among the most widely used types of regularization algorithms, commonly associated with the problem of supervised generalized linear model estimation via penalized maximum likelihood. Its nice properti…

Efficient Elastic Net Regularization for Sparse Linear Models

2015-05-24 · Zachary C. Lipton, Charles Elkan

This paper presents an algorithm for efficient training of sparse linear models with elastic net regularization. Extending previous work on delayed updates, the new algorithm applies stochastic gradient updates to non-ze…

Form

Comparison of 14 different families of classification algorithms on 115 binary datasets

2016-06-02 · Jacques Wainer

We tested 14 very different classification algorithms (random forest, gradient boosting machines, SVM - linear, polynomial, and RBF - 1-hidden-layer neural nets, extreme learning machines, k-nearest neighbors and a baggi…

General ClassificationQuantization

Fast Spatial Memory with Elastic Test-Time Training

2026-04-08 · Ziqiao Ma, Xueyang Yu, Haoyu Zhen, Yuncong Yang 외 arxiv

Large Chunk Test-Time Training (LaCT) has shown strong performance on long-context 3D reconstruction, but its fully plastic inference-time updates remain vulnerable to catastrophic forgetting and overfitting. As a result…

3D Reconstruction

Fast marginal likelihood estimation of penalties for group-adaptive elastic net

2021-01-11 · Mirrelijn M. van Nee, Tim van de Brug, Mark A. van de Wiel

Nowadays, clinical research routinely uses omics data, such as gene expression, for predicting clinical outcomes or selecting markers. Additionally, so-called co-data are often available, providing complementary informat…