paper-with-me

홈 › Papers

Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

2024-08-08 · CVPR 2025 1 · Qirui Jiao, Daoyuan Chen, Yilun Huang, Bolin Ding, Yaliang Li, Ying Shen

High-performance Multimodal Large Language Models (MLLMs) are heavily dependent on data quality. To advance fine-grained image recognition within MLLMs, we introduce a novel data synthesis method inspired by contrastive learning and image difference captioning. Our key idea involves challenging the model to discern both matching and distinct elements by scrutinizing object differences in detailed regions across similar images. We begin by generating pairs of similar images that emphasize object variations. Following this, we employ a Difference Area Generator to pinpoint object differences, and subsequently, a Difference Captions Generator to articulate these differences. This process results in a high-quality dataset of "object replacement" samples, termed Img-Diff, which can be scaled as needed due to its automated nature. We leverage this generated dataset to fine-tune state-of-the-art (SOTA) MLLMs, such as InternVL2, achieving substantial improvements across various image difference and Visual Question Answering tasks. Notably, the trained models significantly outperform existing SOTA models like GPT-4V and Gemini on the MMVP benchmark. Additionally, we conduct comprehensive evaluations to validate the dataset's diversity, quality, and robustness, offering several insights into the synthesis of such contrastive datasets. We release our codes and dataset to encourage further research on multimodal data synthesis and MLLMs' fundamental capabilities for image understanding.

📄 PDF Abstract BibTeX arXiv:2408.04594

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningFine-Grained Image RecognitionObjectQuestion AnsweringVisual Question Answering

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Draw Your Art Dream: Diverse Digital Art Synthesis with Multimodal Guided Diffusion

2022-09-27 · Nisha Huang, Fan Tang, WeiMing Dong, Changsheng Xu

Digital art synthesis is receiving increasing attention in the multimedia community because of engaging the public with art effectively. Current digital art synthesis methods usually use single-modality inputs as guidanc…

Diversity

Discrete Contrastive Diffusion for Cross-Modal Music and Image Generation

2022-06-15 · Ye Zhu, Yu Wu, Kyle Olszewski, Jian Ren 외

Diffusion probabilistic models (DPMs) have become a popular approach to conditional generation, due to their promising results and support for cross-modal synthesis. A key desideratum in conditional synthesis is to achie…

Contrastive LearningDenoisingImage GenerationMusic Generation

Towards General Text-guided Image Synthesis for Customized Multimodal Brain MRI Generation

2024-09-25 · Yulin Wang, Honglin Xiong, Kaicong Sun, Shuwei Bai 외

Multimodal brain magnetic resonance (MR) imaging is indispensable in neuroscience and neurology. However, due to the accessibility of MRI scanners and their lengthy acquisition time, multimodal MR images are not commonly…

Contrastive LearningImage Generation

Contrastive Parameter Disentanglement for Multi-modal Remote Sensing Image Generation

2026-07-26 · Yu Zhang, Wenda Zhao, Haojun Tang, Haipeng Wang arxiv

Existing remote sensing image generation methods are largely confined to single-modality synthesis and therefore fail to exploit the complementary information inherent in multimodal imagery. To address this limitation, w…

Image Generation

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

2026-06-23 · Shikai Qiu, Xiaowen Xu, Benlei Cui, Ting Ma 외 arxiv

General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI safety. We present Yuvion VL, a family of…

Adversarial Robustness