paper-with-me

홈 › Papers

Is Synthetic Data all We Need? Benchmarking the Robustness of Models Trained with Synthetic Images

2024-05-30 · Krishnakant Singh, Thanush Navaratnam, Jannik Holmer, Simone Schaub-Meyer, Stefan Roth

A long-standing challenge in developing machine learning approaches has been the lack of high-quality labeled data. Recently, models trained with purely synthetic data, here termed synthetic clones, generated using large-scale pre-trained diffusion models have shown promising results in overcoming this annotation bottleneck. As these synthetic clone models progress, they are likely to be deployed in challenging real-world settings, yet their suitability remains understudied. Our work addresses this gap by providing the first benchmark for three classes of synthetic clone models, namely supervised, self-supervised, and multi-modal ones, across a range of robustness measures. We show that existing synthetic self-supervised and multi-modal clones are comparable to or outperform state-of-the-art real-image baselines for a range of robustness metrics - shape bias, background bias, calibration, etc. However, we also find that synthetic clones are much more susceptible to adversarial and real-world noise than models trained with real data. To address this, we find that combining both real and synthetic data further increases the robustness, and that the choice of prompt used for generating synthetic images plays an important part in the robustness of synthetic clones.

📄 PDF Abstract BibTeX arXiv:2405.20469

Code (0)

등록된 구현이 없습니다.

Tasks

AllBenchmarking

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study

2024-03-15 · Chenguang Wang, Ruoxi Jia, Xin Liu, Dawn Song

Pre-training image representations from the raw text about images enables zero-shot vision transfer to downstream tasks. Through pre-training on millions of samples collected from the internet, multimodal foundation mode…

Benchmarking

Are Synthetic Corruptions A Reliable Proxy For Real-World Corruptions?

2025-05-07 · Shashank Agnihotri, David Schader, Nico Sharei, Mehmet Ege Kaçar 외

Deep learning (DL) models are widely used in real-world applications but remain vulnerable to distribution shifts, especially due to weather and lighting changes. Collecting diverse real-world data for testing the robust…

BenchmarkingSemantic Segmentation

SynBench: Task-Agnostic Benchmarking of Pretrained Representations using Synthetic Data

2022-10-06 · Ching-Yun Ko, Pin-Yu Chen, Jeet Mohapatra, Payel Das 외

Recent success in fine-tuning large models, that are pretrained on broad data at scale, on downstream tasks has led to a significant paradigm shift in deep learning, from task-centric model design to task-agnostic repres…

BenchmarkingRepresentation Learning

DSLOB: A Synthetic Limit Order Book Dataset for Benchmarking Forecasting Algorithms under Distributional Shift

2022-11-17 · Defu Cao, Yousef El-Laham, Loc Trinh, Svitlana Vyetrenko 외

In electronic trading markets, limit order books (LOBs) provide information about pending buy/sell orders at various price levels for a given security. Recently, there has been a growing interest in using LOB data for re…

BenchmarkingTime SeriesTime Series Analysis

WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attribution

2025-04-28 · Pietro Bongini, Sara Mandelli, Andrea Montibeller, Mirko Casu 외

Synthetic image source attribution is an open challenge, with an increasing number of image generators being released yearly. The complexity and the sheer number of available generative techniques, as well as the scarcit…

BenchmarkingImage AttributionSynthetic Image Attribution