paper-with-me

Papers

ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans

2025-07-30 · Ananya Sadana, Yash Kumar Lal, Jiawei Zhou arxiv

Understanding causal relationships across modalities is a core challenge for multimodal models operating in real-world environments. We introduce ISO-Bench, a benchmark for evaluating whether models can infer causal dependencies between visual observations and procedural text. Each example presents an image of a task step and a text snippet from a plan, with the goal of deciding whether the visual step occurs before or after the referenced text step. Evaluation results on ten frontier vision-language models show underwhelming performance: the best zero-shot F1 is only 0.57, and chain-of-thought reasoning yields only modest gains (up to 0.62 F1), largely behind humans (0.98 F1). Our analysis further highlights concrete directions for improving causal understanding in multimodal models.

📄 PDF Abstract BibTeX arXiv:2507.23135

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LogicGaze: Benchmarking Causal Consistency in Visual Narratives via Counterfactual Verification

2026-01-30 · Rory Driscoll, Alexandros Christoforos, Chadbourne Davis arxiv

While sequential reasoning enhances the capability of Vision-Language Models (VLMs) to execute complex multimodal tasks, their reliability in grounding these reasoning chains within actual visual evidence remains insuffi…

Multimodal Reasoning

Multimodal Causal Reasoning Benchmark: Challenging Vision Large Language Models to Infer Causal Links Between Siamese Images

2024-08-15 · Zhiyuan Li, Heng Wang, Dongnan Liu, Chaoyi Zhang 외

Large Language Models (LLMs) have showcased exceptional ability in causal reasoning from textual information. However, will these causalities remain straightforward for Vision Large Language Models (VLLMs) when only visu…

Image GenerationSentence

Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing

2025-04-03 · Xiangyu Zhao, Peiyuan Zhang, Kexian Tang, Hao Li 외

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preservi…

BenchmarkingLogical Reasoning

HourVideo: 1-Hour Video-Language Understanding

2024-11-07 · Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota 외

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, tempora…

BenchmarkingcounterfactualMultiple-choiceRetrieval+1

M$^3$Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions

2026-06-05 · Zhengjun Huang, Wenxuan Liu, Zhoujin Tian, Wei Chen 외 arxiv

Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visuals and straightforward content, evaluating neither reasoning over aut…