Improving short text classification through global augmentation methods
We study the effect of different approaches to text augmentation. To do this we use 3 datasets that include social media and formal text in the form of news articles. Our goal is to provide insights for practitioners and researchers on making choices for augmentation for classification use cases. We observe that Word2vec-based augmentation is a viable option when one does not have access to a formal synonym model (like WordNet-based augmentation). The use of \emph{mixup} further improves performance of all text based augmentations and reduces the effects of overfitting on a tested deep learning model. Round-trip translation with a translation service proves to be harder to use due to cost and as such is less accessible for both normal and low resource use-cases.
Code (1)
Tasks
ArticlesClassificationGeneral ClassificationText Augmentationtext-classificationText ClassificationTranslationSimilar Papers 제목 키워드 기반
Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts
Sentiment classification in short text datasets faces significant challenges such as class imbalance, limited training samples, and the inherent subjectivity of sentiment labels -- issues that are further intensified by …
Sentiment AnalysisData AugmentationA Simple Graph Contrastive Learning Framework for Short Text Classification
Short text classification has gained significant attention in the information age due to its prevalence and real-world applications. Recent advancements in graph learning combined with contrastive learning have shown pro…
Contrastive LearningData AugmentationGraph Learningtext-classification+1Low resource language dataset creation, curation and classification: Setswana and Sepedi -- Extended Abstract
The recent advances in Natural Language Processing have only been a boon for well represented languages, negating research in lesser known global languages. This is in part due to the availability of curated data and res…
Data AugmentationGeneral ClassificationTopic ClassificationThe Effectiveness of Data Augmentation in Image Classification using Deep Learning
In this paper, we explore and compare multiple solutions to the problem of data augmentation in image classification. Previous work has demonstrated the effectiveness of data augmentation through simple techniques, such …
Data AugmentationGeneral Classificationimage-classificationImage ClassificationFrom Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
Pretrained vision-language models (VLMs), e.g., CLIP, demonstrate impressive zero-shot capabilities on downstream tasks. Prior research highlights the crucial role of visual augmentation techniques, like random cropping,…
feature selectionOut-of-Distribution GeneralizationTest-time Adaptationzero-shot-classification+1