paper-with-me

Papers

Exploring the Potential of Synthetic Data to Replace Real Data

2024-08-26 · Hyungtae Lee, Yan Zhang, Heesung Kwon, Shuvra S. Bhattacharrya

The potential of synthetic data to replace real data creates a huge demand for synthetic data in data-hungry AI. This potential is even greater when synthetic data is used for training along with a small number of real images from domains other than the test domain. We find that this potential varies depending on (i) the number of cross-domain real images and (ii) the test set on which the trained model is evaluated. We introduce two new metrics, the train2test distance and $\text{AP}_\text{t2t}$, to evaluate the ability of a cross-domain training set using synthetic data to represent the characteristics of test instances in relation to training performance. Using these metrics, we delve deeper into the factors that influence the potential of synthetic data and uncover some interesting dynamics about how synthetic data impacts training performance. We hope these discoveries will encourage more widespread use of synthetic data.

📄 PDF Abstract BibTeX arXiv:2408.14559

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Exploring the Potential of AI-Generated Synthetic Datasets: A Case Study on Telematics Data with ChatGPT

2023-06-23 · Ryan Lingo

This research delves into the construction and utilization of synthetic datasets, specifically within the telematics sphere, leveraging OpenAI's powerful language model, ChatGPT. Synthetic datasets present an effective s…

DescriptiveLanguage Modelling

AutoSynth: Learning to Generate 3D Training Data for Object Point Cloud Registration

2023-09-20 · ICCV 2023 1 · Zheng Dang, Mathieu Salzmann

In the current deep learning paradigm, the amount and quality of training data are as critical as the network architecture and its training details. However, collecting, processing, and annotating real data at scale is d…

Meta-LearningPoint Cloud Registration

SimVQA: Exploring Simulated Environments for Visual Question Answering

2022-03-31 · CVPR 2022 1 · Paola Cascante-Bonilla, Hui Wu, Letao Wang, Rogerio Feris 외

Existing work on VQA explores data augmentation to achieve better generalization by perturbing the images in the dataset or modifying the existing questions and answers. While these methods exhibit good performance, the …

Data AugmentationDiversityQuestion AnsweringVisual Question Answering+1

Can Synthetic Translations Improve Bitext Quality?

2021-10-16 · ACL ARR October 2021 10 · Anonymous

Synthetic translations have been used for a wide range of NLP tasks primarily as a means of data augmentation. This work explores instead, how we can use synthetic translations to selectively replace potentially imperfec…

Data AugmentationNMT

Synthetic Data for Model Selection

2021-05-03 · Alon Shoshan, Nadav Bhonker, Igor Kviatkovsky, Matan Fintz 외

Recent breakthroughs in synthetic data generation approaches made it possible to produce highly photorealistic images which are hardly distinguishable from real ones. Furthermore, synthetic generation pipelines have the …

image-classificationImage ClassificationmodelModel Selection+1