paper-with-me

Papers

Source Code Data Augmentation for Deep Learning: A Survey

2023-05-31 · Terry Yue Zhuo, Zhou Yang, Zhensu Sun, YuFei Wang, Li Li, Xiaoning Du, Zhenchang Xing, David Lo

The increasingly popular adoption of deep learning models in many critical source code tasks motivates the development of data augmentation (DA) techniques to enhance training data and improve various capabilities (e.g., robustness and generalizability) of these models. Although a series of DA methods have been proposed and tailored for source code models, there lacks a comprehensive survey and examination to understand their effectiveness and implications. This paper fills this gap by conducting a comprehensive and integrative survey of data augmentation for source code, wherein we systematically compile and encapsulate existing literature to provide a comprehensive overview of the field. We start with an introduction of data augmentation in source code and then provide a discussion on major representative approaches. Next, we highlight the general strategies and techniques to optimize the DA quality. Subsequently, we underscore techniques useful in real-world source code scenarios and downstream tasks. Finally, we outline the prevailing challenges and potential opportunities for future research. In essence, we aim to demystify the corpus of existing literature on source code DA for deep learning, and foster further exploration in this sphere. Complementing this, we present a continually updated GitHub repository that hosts a list of update-to-date papers on DA for source code modeling, accessible at \url{https://github.com/terryyz/DataAug4Code}.

📄 PDF Abstract BibTeX arXiv:2305.19915

Code (1)

terryyz/dataaug4code 공식 구현

Tasks

Data AugmentationDeep LearningSurvey

Similar Papers 제목 키워드 기반

Web Table Extraction, Retrieval and Augmentation: A Survey

2020-02-01 · Shuo Zhang, Krisztian Balog

Tables are a powerful and popular tool for organizing and manipulating data. A vast number of tables can be found on the Web, which represents a valuable knowledge resource. The objective of this survey is to synthesize …

Question AnsweringRetrievalSurveyTable Extraction+1

Bridging the Linguistic Divide: A Survey on Leveraging Large Language Models for Machine Translation

2025-04-02 · Baban Gain, Dibyanayan Bandyopadhyay, Asif Ekbal

The advent of Large Language Models (LLMs) has significantly reshaped the landscape of machine translation (MT), particularly for low-resource languages and domains that lack sufficient parallel corpora, linguistic tools…

Cross-Lingual TransferDecoderMachine Translationparameter-efficient fine-tuning+3

An Empirical Survey of Data Augmentation for Limited Data Learning in NLP

2021-06-14 · Jiaao Chen, Derek Tam, Colin Raffel, Mohit Bansal 외

NLP has achieved great progress in the past decade through the use of neural models and large labeled datasets. The dependence on abundant data prevents NLP models from being applied to low-resource settings or novel tas…

Data AugmentationNews ClassificationSentence

An Empirical Survey of Data Augmentation \\for Limited Data Learning in NLP

2021-08-17 · ACL ARR August 2021 8 · Anonymous

NLP has achieved great progress in the past decade through the use of neural models and large labeled datasets. The dependence on abundant data prevents NLP models from being applied to low-resource settings or novel ta…

Data AugmentationNews ClassificationSentence

Image Data Augmentation Approaches: A Comprehensive Survey and Future directions

2023-01-07 · Teerath Kumar, Alessandra Mileo, Rob Brennan, Malika Bendechache

Deep learning (DL) algorithms have shown significant performance in various computer vision tasks. However, having limited labelled data lead to a network overfitting problem, where network performance is bad on unseen d…

Data Augmentationimage-classificationImage Classificationobject-detection+3