paper-with-me

홈 › Papers

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?

2026-02-03 · Jingyi Zhang, Tianyi Lin, Huanjin Yao, Xiang Lan, Shunyu Liu, Jiaxing Huang arxiv

In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks. To this end, we propose Collective Adversarial Data Synthesis (CADS), a novel and general approach to synthesize high-quality, diverse and challenging multimodal data for MLLMs. The core idea of CADS is to leverage collective intelligence to ensure high-quality and diverse generation, while exploring adversarial learning to synthesize challenging samples for effectively driving model improvement. Specifically, CADS operates with two cyclic phases, i.e., Collective Adversarial Data Generation (CAD-Generate) and Collective Adversarial Data Judgment (CAD-Judge). CAD-Generate leverages collective knowledge to jointly generate new and diverse multimodal data, while CAD-Judge collaboratively assesses the quality of synthesized data. In addition, CADS introduces an Adversarial Context Optimization mechanism to optimize the generation context to encourage challenging and high-value data generation. With CADS, we construct MMSynthetic-20K and train our model R1-SyntheticVL, which demonstrates superior performance on various benchmarks.

📄 PDF Abstract BibTeX arXiv:2602.03300

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D Scenes

2025-10-30 · Yukun Huang, Jiwen Yu, Yanning Zhou, Jianan Wang 외 arxiv

There are two prevalent ways to constructing 3D scenes: procedural generation and 2D lifting. Among them, panorama-based 2D lifting has emerged as a promising technique, leveraging powerful 2D generative priors to produc…

Scene Generation

SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision Encoders

2026-06-07 · Seunghyun Lee, Byoungkwon Kim, Jaehyun Nam, Kyungmin Lee 외 arxiv

The rapid advancement of generative models has blurred the boundary between synthetic and real imagery, creating an urgent need for reliable deepfake detection. Yet most existing approaches rely on massive real--fake dat…

DeepFake Detection

Is synthetic data from generative models ready for image recognition?

2022-10-14 · Ruifei He, Shuyang Sun, Xin Yu, Chuhui Xue 외

Recent text-to-image generation models have shown promising results in generating high-fidelity photo-realistic images. Though the results are astonishing to human eyes, how applicable these generated images are for reco…

Image GenerationText to Image GenerationText-to-Image GenerationTransfer Learning

HaDR: Applying Domain Randomization for Generating Synthetic Multimodal Dataset for Hand Instance Segmentation in Cluttered Industrial Environments

2023-04-12 · Stefan Grushko, Aleš Vysocký, Jakub Chlebek, Petr Prokop

This study uses domain randomization to generate a synthetic RGB-D dataset for training multimodal instance segmentation models, aiming to achieve colour-agnostic hand localization in cluttered industrial environments. D…

Hand DetectionInstance SegmentationSemantic Segmentation

The Synthetic Media Shift: Tracking the Rise, Virality, and Detectability of AI-Generated Multimodal Misinformation

2026-04-15 · Zacharias Chrysidis, Stefanos-Iordanis Papadopoulos, Symeon Papadopoulos arxiv

As generative AI advances, the distinction between authentic and synthetic media is increasingly blurred, challenging the integrity of online information. In this study, we present CONVEX, a large-scale dataset of multim…