paper-with-me

홈 › Papers

How Effective is Task-Agnostic Data Augmentation for Pretrained Transformers?

2020-10-05 · Findings of the Association for Computational Linguistics 2020 · Shayne Longpre, Yu Wang, Christopher DuBois

Task-agnostic forms of data augmentation have proven widely effective in computer vision, even on pretrained models. In NLP similar results are reported most commonly for low data regimes, non-pretrained models, or situationally for pretrained models. In this paper we ask how effective these techniques really are when applied to pretrained transformers. Using two popular varieties of task-agnostic data augmentation (not tailored to any particular task), Easy Data Augmentation (Wei and Zou, 2019) and Back-Translation (Sennrichet al., 2015), we conduct a systematic examination of their effects across 5 classification tasks, 6 datasets, and 3 variants of modern pretrained transformers, including BERT, XLNet, and RoBERTa. We observe a negative result, finding that techniques which previously reported strong improvements for non-pretrained models fail to consistently improve performance for pretrained transformers, even when training data is limited. We hope this empirical analysis helps inform practitioners where data augmentation techniques may confer improvements.

📄 PDF Abstract BibTeX arXiv:2010.01764

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationTranslation

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
WordPiece 설명 없음
Multi-Head Attention 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.

Similar Papers 제목 키워드 기반

RepAugment: Input-Agnostic Representation-Level Augmentation for Respiratory Sound Classification

2024-05-05 · June-Woo Kim, Miika Toikkanen, Sangmin Bae, Minseok Kim 외

Recent advancements in AI have democratized its deployment as a healthcare assistant. While pretrained models from large-scale visual and audio datasets have demonstrably generalized to this task, surprisingly, no studie…

Data AugmentationSound Classification

TAP-CT: 3D Task-Agnostic Pretraining of Computed Tomography Foundation Models

2025-11-30 · Tim Veenboer, George Yiasemis, Eric Marcus, Vivien Van Veldhuizen 외 arxiv

Existing foundation models (FMs) in the medical domain often require extensive fine-tuning or rely on training resource-intensive decoders, while many existing encoders are pretrained with objectives biased toward specif…

The Efficacy of Semantics-Preserving Transformations in Self-Supervised Learning for Medical Ultrasound

2025-04-10 · Blake VanBerlo, Alexander Wong, Jesse Hoey, Robert Arntfield

Data augmentation is a central component of joint embedding self-supervised learning (SSL). Approaches that work for natural images may not always be effective in medical imaging tasks. This study systematically investig…

ClassificationData AugmentationDiagnosticLine Detection+1

Learning Data Augmentation Schedules for Natural Language Processing

2021-11-01 · EMNLP (insights) 2021 11 · Daphné Chopard, Matthias S. Treder, Irena Spasić

Despite its proven efficiency in other fields, data augmentation is less popular in the context of natural language processing (NLP) due to its complexity and limited results. A recent study (Longpre et al., 2020) showed…

Data AugmentationScheduling

Transfer Learning with EfficientNet for Accurate Leukemia Cell Classification

2025-08-04 · Faisal Ahmed arxiv

Accurate classification of Acute Lymphoblastic Leukemia (ALL) from peripheral blood smear images is essential for early diagnosis and effective treatment planning. This study investigates the use of transfer learning wit…

Transfer LearningData Augmentation