Roles of Words: What Should (n’t) Be Augmented in Text Augmentation on Text Classification Tasks?
Text augmentation techniques are widely used in text classification problems to improve the performance of classifiers, especially in low-resource scenarios. Previous text-editing-based methods augment the text in a non-selective manner: the words in the text are treated without difference during augmentation, which may result in unsatisfactory augmented samples. In this work, we present four kinds of roles of words (ROWs) which have different functions in text classification tasks, and design effective methods to automatically extract these ROWs based on statistical and semantic perspectives. Systematic experiments are conducted on what ROWs should (n't) be augmented during augmentation for classification tasks. Based on these experiments, we discover some interesting and instructive potential patterns that certain ROWs are especially suitable or unsuitable for certain augmentation operations. Guided by these patterns, we propose a set of Selective Text Augmentation (STA) operations, which significantly outperform traditional methods and show outstanding generalization performance.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationText Augmentationtext-classificationText ClassificationSimilar Papers 제목 키워드 기반
What Have Been Learned & What Should Be Learned? An Empirical Study of How to Selectively Augment Text for Classification
Text augmentation techniques are widely used in text classification problems to improve the performance of classifiers, especially in low-resource scenarios. Whilst lots of creative text augmentation methods have been de…
ClassificationText Augmentationtext-classificationText ClassificationSelective Text Augmentation with Word Roles for Low-Resource Text Classification
Data augmentation techniques are widely used in text classification tasks to improve the performance of classifiers, especially in low-resource scenarios. Most previous methods conduct text augmentation without consideri…
ClassificationData AugmentationLanguage ModellingLarge Language Model+5MulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval
Training-free zero-shot composed image retrieval finds a target image in a gallery from a reference image and a text edit without learning from task-specific image triplets. Existing methods typically describe the target…
Image RetrievalUnsupervised Learning of Entailment-Vector Word Embeddings
Entailment vectors are a principled way to encode in a vector what information is known and what is unknown. They are designed to model relations where one vector should include all the information in another vector, ca…
Word EmbeddingsJoint Learning Templates and Slots for Event Schema Induction
Automatic event schema induction (AESI) means to extract meta-event from raw text, in other words, to find out what types (templates) of event may exist in the raw text and what roles (slots) may exist in each event type…
Image SegmentationSemantic SegmentationSentence