paper-with-me

홈 › Papers

Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities

2025-01-31 · Yaping Chai, Haoran Xie, Joe S. Qin

The increasing size and complexity of pre-trained language models have demonstrated superior performance in many applications, but they usually require large training datasets to be adequately trained. Insufficient training sets could unexpectedly make the model overfit and fail to cope with complex tasks. Large language models (LLMs) trained on extensive corpora have prominent text generation capabilities, which improve the quality and quantity of data and play a crucial role in data augmentation. Specifically, distinctive prompt templates are given in personalised tasks to guide LLMs in generating the required content. Recent promising retrieval-based techniques further improve the expressive performance of LLMs in data augmentation by introducing external knowledge to enable them to produce more grounded-truth data. This survey provides an in-depth analysis of data augmentation in LLMs, classifying the techniques into Simple Augmentation, Prompt-based Augmentation, Retrieval-based Augmentation and Hybrid Augmentation. We summarise the post-processing approaches in data augmentation, which contributes significantly to refining the augmented data and enabling the model to filter out unfaithful content. Then, we provide the common tasks and evaluation metrics. Finally, we introduce existing challenges and future opportunities that could bring further improvement to data augmentation.

📄 PDF Abstract BibTeX arXiv:2501.18845

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationRetrievalText Generation

Similar Papers 제목 키워드 기반

Text Data Augmentation: Towards better detection of spear-phishing emails

2020-07-04 · Mehdi Regina, Maxime Meyer, Sébastien Goutal

Text data augmentation, i.e., the creation of new textual data from an existing text, is challenging. Indeed, augmentation transformations should take into account language complexity while being relevant to the target N…

Data AugmentationGeneral ClassificationLanguage ModelingLanguage Modelling+6

Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges

2024-03-05 · Bosheng Ding, Chengwei Qin, Ruochen Zhao, Tianze Luo 외

In the rapidly evolving field of large language models (LLMs), data augmentation (DA) has emerged as a pivotal technique for enhancing model performance by diversifying training examples without the need for additional d…

Data AugmentationSurvey

LaMP: When Large Language Models Meet Personalization

2023-04-22 · Alireza Salemi, Sheshera Mysore, Michael Bendersky, Hamed Zamani

This paper highlights the importance of personalization in large language models and introduces the LaMP benchmark -- a novel benchmark for training and evaluating language models for producing personalized outputs. LaMP…

Language ModelingLanguage ModellingNatural Language UnderstandingRetrieval+3

Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning

2023-05-21 · Qiming Bao, Alex Yuxuan Peng, Zhenyun Deng, Wanjun Zhong 외

Combining large language models with logical reasoning enhances their capacity to address problems in a robust and reliable manner. Nevertheless, the intricate nature of logical reasoning poses challenges when gathering …

Abstract Meaning RepresentationContrastive LearningData AugmentationLanguage Modelling+11

A Survey on Data Augmentation in Large Model Era

2024-01-27 · Yue Zhou, Chenlu Guo, Xu Wang, Yi Chang 외

Large models, encompassing large language and diffusion models, have shown exceptional promise in approximating human-level intelligence, garnering significant interest from both academic and industrial spheres. However,…

Audio Signal ProcessingData AugmentationImage AugmentationSurvey+1