paper-with-me

홈 › Papers

Comparing effectiveness of regularization methods on text classification: Simple and complex model in data shortage situation

2024-02-27 · Jongga Lee, Jaeseung Yim, Seohee Park, Changwon Lim

Text classification is the task of assigning a document to a predefined class. However, it is expensive to acquire enough labeled documents or to label them. In this paper, we study the regularization methods' effects on various classification models when only a few labeled data are available. We compare a simple word embedding-based model, which is simple but effective, with complex models (CNN and BiLSTM). In supervised learning, adversarial training can further regularize the model. When an unlabeled dataset is available, we can regularize the model using semi-supervised learning methods such as the Pi model and virtual adversarial training. We evaluate the regularization effects on four text classification datasets (AG news, DBpedia, Yahoo! Answers, Yelp Polarity), using only 0.1% to 0.5% of the original labeled training documents. The simple model performs relatively well in fully supervised learning, but with the help of adversarial training and semi-supervised learning, both simple and complex models can be regularized, showing better results for complex models. Although the simple model is robust to overfitting, a complex model with well-designed prior beliefs can be also robust to overfitting.

📄 PDF Abstract BibTeX arXiv:2403.00825

Code (0)

등록된 구현이 없습니다.

Tasks

text-classificationText Classification

Similar Papers 제목 키워드 기반

Multiview Hessian Regularization for Image Annotation

2019-04-23 · Weifeng Liu, DaCheng Tao

The rapid development of computer hardware and Internet technology makes large scale data dependent models computationally tractable, and opens a bright avenue for annotating images through innovative machine learning al…

General Classification

Epsilon Consistent Mixup: Structural Regularization with an Adaptive Consistency-Interpolation Tradeoff

2021-04-19 · Vincent Pisztora, Yanglan Ou, Xiaolei Huang, Francesca Chiaromonte 외

In this paper we propose $\epsilon$-Consistent Mixup ($\epsilon$mu). $\epsilon$mu is a data-based structural regularization technique that combines Mixup's linear interpolation with consistency regularization in the Mixu…

Revisiting Orthogonality Regularization: A Study for Convolutional Neural Networks in Image Classification

2022-06-23 · IEEE Access 2022 6 · Taehyeon Kim, Se-Young Yun

Recent research in deep Convolutional Neural Networks(CNN) faces the challenges of vanishing/exploding gradient issues, training instability, and feature redundancy. Orthogonality Regularization(OR), which introduces a p…

image-classificationImage Classification

Manifold Based Low-rank Regularization for Image Restoration and Semi-supervised Learning

2017-02-09 · Rongjie Lai, Jia Li

Low-rank structures play important role in recent advances of many problems in image science and data science. As a natural extension of low-rank structures for data with nonlinear structures, the concept of the low-dime…

Image InpaintingImage ReconstructionImage RestorationImage Super-Resolution+1

Robust Multi-class Feature Selection via $l_{2,0}$-Norm Regularization Minimization

2020-10-08 · Zhenzhen Sun, Yuanlong Yu

Feature selection is an important data pre-processing in data mining and machine learning, which can reduce feature size without deteriorating model's performance. Recently, sparse regression based feature selection meth…

feature selection