paper-with-me

Papers

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts

2026-04-30 · Bingli Wang, Huanze Tang, Haijun Lv, Zhishan Lin, Lixin Gu, Lei Feng, Qipeng Guo, Kai Chen arxiv

In recent years, Multimodal Large Language Models (MLLMs) have achieved remarkable progress on a wide range of multimodal benchmarks. Despite these advances, most existing benchmarks mainly focus on single-image or multi-image comprehension. In real-world scenarios such as document reading, information is often presented as interleaved multimodel contexts. This requires MLLMs not only to recognize the content of individual images, but also to identify relevant textual and visual evidence, establish fine-grained alignments between them, and reason over these aligned signals in interleaved contexts based on contextual evidence. However, there is still a lack of systematic benchmarks for quantifying the fine-grained understanding ability of MLLMs in interleaved image-text contexts. To fill this gap, we propose COHERENCE, a benchmark designed to evaluate the ability of MLLMs to recover fine-grained image-text correspondences in interleaved multimodal contexts. COHERENCE covers interleaved image-text content from four representative domains and contains 6,161 high-quality questions. Moreover, we perform a six-type error analysis, enabling fine-grained attribution of failures in interleaved image-text understanding to the specific capabilities missing in current MLLMs.

📄 PDF Abstract BibTeX arXiv:2604.27389

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation

2025-02-03 · Haibo Tong, Zhaoyang Wang, Zhaorun Chen, Haonian Ji 외

Recent advancements in video generation have significantly improved the ability to synthesize videos from text instructions. However, existing models still struggle with key challenges such as instruction misalignment, c…

BenchmarkingFairnessHallucinationMixture-of-Experts+1

SNaC: Coherence Error Detection for Narrative Summarization

2022-05-19 · Tanya Goyal, Junyi Jessy Li, Greg Durrett

Progress in summarizing long texts is inhibited by the lack of appropriate evaluation frameworks. When a long summary must be produced to appropriately cover the facets of that text, that summary needs to present a coher…

BenchmarkingCoherence EvaluationDocument Summarization

FarSLIP: Discovering Effective CLIP Adaptation for Fine-Grained Remote Sensing Understanding

2025-11-18 · Zhenshi Li, Weikang Yu, Dilxat Muhtar, Xueliang Zhang 외 arxiv

As CLIP's global alignment limits its ability to capture fine-grained details, recent efforts have focused on enhancing its region-text alignment. However, current remote sensing (RS)-specific CLIP variants still inherit…

Semantic SegmentationText Retrieval

ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction

2025-06-10 · Juan Yeo, Soonwoo Cha, Jiwoo Song, Hyunbin Jin 외

Vision-language models such as CLIP have recently propelled open-vocabulary dense prediction tasks by enabling recognition of a broad range of visual concepts. However, CLIP still struggles with fine-grained, region-leve…

object-detectionObject DetectionOpen-vocabulary object detectionOpen Vocabulary Object Detection+1

African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification

2024-06-20 · Gregor Geigle, Radu Timofte, Goran Glavaš

Recent Large Vision-Language Models (LVLMs) demonstrate impressive abilities on numerous image understanding and reasoning tasks. The task of fine-grained object classification (e.g., distinction between \textit{animal s…

BenchmarkingClassificationMultiple-choiceObject+1