paper-with-me

홈 › Papers

VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models

2024-06-14 · Chenyu Zhou, Mengdan Zhang, Peixian Chen, Chaoyou Fu, Yunhang Shen, Xiawu Zheng, Xing Sun, Rongrong Ji

The swift progress of Multi-modal Large Models (MLLMs) has showcased their impressive ability to tackle tasks blending vision and language. Yet, most current models and benchmarks cater to scenarios with a narrow scope of visual and textual contexts. These models often fall short when faced with complex comprehension tasks, which involve navigating through a plethora of irrelevant and potentially misleading information in both text and image forms. To bridge this gap, we introduce a new, more demanding task known as Interleaved Image-Text Comprehension (IITC). This task challenges models to discern and disregard superfluous elements in both images and text to accurately answer questions and to follow intricate instructions to pinpoint the relevant image. In support of this task, we further craft a new VEGA dataset, tailored for the IITC task on scientific content, and devised a subtask, Image-Text Association (ITA), to refine image-text correlation skills. Our evaluation of four leading closed-source models, as well as various open-source models using VEGA, underscores the rigorous nature of IITC. Even the most advanced models, such as Gemini-1.5-pro and GPT4V, only achieved modest success. By employing a multi-task, multi-scale post-training strategy, we have set a robust baseline for MLLMs on the IITC task, attaining an $85.8\%$ accuracy rate in image association and a $0.508$ Rouge score. These results validate the effectiveness of our dataset in improving MLLMs capabilities for nuanced image-text comprehension.

📄 PDF Abstract BibTeX arXiv:2406.10228

Code (0)

등록된 구현이 없습니다.

Tasks

Reading Comprehension

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
VEGA 설명 없음

Similar Papers 제목 키워드 기반

WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation

2025-11-14 · Wei Chow, Jiachun Pan, Yongyuan Liang, Mingze Zhou 외 arxiv

Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primarily on single-turn interactions, failing …

Image GenerationImage Editing

MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models

2024-10-14 · Peng Xia, Siwei Han, Shi Qiu, Yiyang Zhou 외

Interleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal learning. Despite significant advancemen…

Multiple-choice

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

2024-07-03 · Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao 외

We present InternLM-XComposer-2.5 (IXC-2.5), a versatile large-vision language model that supports long-contextual input and output. IXC-2.5 excels in various text-image comprehension and composition applications, achiev…

ArticlesImage ComprehensionLanguage ModelingLanguage Modelling+4

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts

2026-04-30 · Bingli Wang, Huanze Tang, Haijun Lv, Zhishan Lin 외 arxiv

In recent years, Multimodal Large Language Models (MLLMs) have achieved remarkable progress on a wide range of multimodal benchmarks. Despite these advances, most existing benchmarks mainly focus on single-image or multi…

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

2023-09-26 · Pan Zhang, Xiaoyi Dong, Bin Wang, Yuhang Cao 외

We propose InternLM-XComposer, a vision-language large model that enables advanced image-text comprehension and composition. The innovative nature of our model is highlighted by three appealing properties: 1) Interleaved…

ArticlesImage ComprehensionMMEReading Comprehension+1