paper-with-me

홈 › Papers

Learning Data Augmentation Schedules for Natural Language Processing

2021-11-01 · EMNLP (insights) 2021 11 · Daphné Chopard, Matthias S. Treder, Irena Spasić

Despite its proven efficiency in other fields, data augmentation is less popular in the context of natural language processing (NLP) due to its complexity and limited results. A recent study (Longpre et al., 2020) showed for example that task-agnostic data augmentations fail to consistently boost the performance of pretrained transformers even in low data regimes. In this paper, we investigate whether data-driven augmentation scheduling and the integration of a wider set of transformations can lead to improved performance where fixed and limited policies were unsuccessful. Our results suggest that, while this approach can help the training process in some settings, the improvements are unsubstantial. This negative result is meant to help researchers better understand the limitations of data augmentation for NLP.

📄 PDF Abstract BibTeX

Code (1)

chopardda/ldas-nlp 공식 구현 tf

Tasks

Data AugmentationScheduling

Similar Papers 제목 키워드 기반

Population Based Augmentation: Efficient Learning of Augmentation Policy Schedules

2019-05-14 · Daniel Ho, Eric Liang, Ion Stoica, Pieter Abbeel 외

A key challenge in leveraging data augmentation for neural network training is choosing an effective augmentation policy from a large search space of candidate operations. Properly chosen augmentation policies can lead t…

Data AugmentationImage Augmentation

LINDA: Unsupervised Learning to Interpolate in Natural Language Processing

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Despite the success of mixup in data augmentation, its applicability to natural language processing (NLP) tasks has been limited due to the discrete and variable-length nature of natural languages. Recent studies have th…

Data Augmentationtext-classificationText Classification

LINDA: Unsupervised Learning to Interpolate in Natural Language Processing

2021-12-28 · Yekyung Kim, Seohyeong Jeong, Kyunghyun Cho

Despite the success of mixup in data augmentation, its applicability to natural language processing (NLP) tasks has been limited due to the discrete and variable-length nature of natural languages. Recent studies have th…

Data Augmentationtext-classificationText Classification

Investigating Masking-based Data Generation in Language Models

2023-06-16 · Ed S. Ma

The current era of natural language processing (NLP) has been defined by the prominence of pre-trained language models since the advent of BERT. A feature of BERT and models with similar architecture is the objective of …

Data AugmentationLanguage ModelingLanguage ModellingMasked Language Modeling

How Data Augmentation affects Optimization for Linear Regression

2020-10-21 · NeurIPS 2021 12 · Boris Hanin, Yi Sun

Though data augmentation has rapidly emerged as a key tool for optimization in modern machine learning, a clear picture of how augmentation schedules affect optimization and interact with optimization hyperparameters suc…

Data AugmentationregressionStochastic Optimization