paper-with-me

홈 › Papers

VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors

2025-12-07 · Wenbo Lyu, Yingjun Du, Jinglin Zhao, Xianton Zhen, Ling Shao arxiv

Understanding multi-image, multi-turn scenarios is a critical yet underexplored capability for Large Vision-Language Models (LVLMs). Existing benchmarks predominantly focus on static or horizontal comparisons -- e.g., spotting visual differences or assessing appropriateness -- while relying heavily on language cues. Such settings overlook progressive, context-dependent reasoning and the challenge of visual-to-visual inference. To bridge this gap, we present VisChainBench, a large-scale benchmark designed to rigorously evaluate LVLMs' ability to perform multi-step visual reasoning across sequential, interdependent tasks with minimal language guidance. VisChainBench contains 1,457 tasks spanning over 20,000 images across three diverse domains (e.g., daily scenarios, engineering troubleshooting), structured to mimic real-world decision-making processes. Uniquely, the benchmark is constructed using a multi-agent generation pipeline, ensuring high visual diversity and controlled language bias. All the benchmark data and code for benchmark construction are available for viewing and download via following Link: https://huggingface.co/datasets/eyehole/VisChainBench

📄 PDF Abstract BibTeX arXiv:2512.06759

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

TheaterGen: Character Management with LLM for Consistent Multi-turn Image Generation

2024-04-29 · Junhao Cheng, Baiqiao Yin, Kaixin Cai, Minbin Huang 외

Recent advances in diffusion models can generate high-quality and stunning images from text. However, multi-turn image generation, which is of high demand in real-world scenarios, still faces challenges in maintaining se…

DenoisingImage GenerationManagementStory Generation

MMCR: Advancing Visual Language Model in Multimodal Multi-Turn Contextual Reasoning

2025-03-24 · Dawei Yan, Yang Li, Qing-Guo Chen, Weihua Luo 외

Compared to single-turn dialogue, multi-turn dialogue involving multiple images better aligns with the needs of real-world human-AI interactions. Additionally, as training data, it provides richer contextual reasoning in…

DiagnosticLanguage ModelingLanguage ModellingPrompt Engineering

CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark

2025-10-30 · Jiaqi Wang, Xiao Yang, Kai Sun, Parth Suresh 외 arxiv

Wearable devices such as smart glasses are transforming the way people interact with their surroundings, enabling users to seek information regarding entities in their view. Multi-Modal Retrieval-Augmented Generation (MM…

MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

2024-06-17 · Ziyu Liu, Tao Chu, Yuhang Zang, Xilin Wei 외

Generating natural and meaningful responses to communicate with multi-modal human inputs is a fundamental capability of Large Vision-Language Models(LVLMs). While current open-source LVLMs demonstrate promising performan…

Visual Question Answering

The Neural Painter: Multi-Turn Image Generation

2018-06-16 · Ryan Y. Benmalek, Claire Cardie, Serge Belongie, Xiadong He 외

In this work we combine two research threads from Vision/ Graphics and Natural Language Processing to formulate an image generation task conditioned on attributes in a multi-turn setting. By multiturn, we mean the image …

BenchmarkingConditional Image GenerationImage Generation