paper-with-me

Papers

LogicGaze: Benchmarking Causal Consistency in Visual Narratives via Counterfactual Verification

2026-01-30 · Rory Driscoll, Alexandros Christoforos, Chadbourne Davis arxiv

While sequential reasoning enhances the capability of Vision-Language Models (VLMs) to execute complex multimodal tasks, their reliability in grounding these reasoning chains within actual visual evidence remains insufficiently explored. We introduce LogicGaze, a novel benchmark framework designed to rigorously interrogate whether VLMs can validate sequential causal chains against visual inputs, specifically targeting the pervasive issue of hallucination. Curated from 40,000 video segments from ShareGPT4Video and a subset of Flickr30k imagery, LogicGaze integrates causal sequences with visually contradictory yet linguistically plausible perturbations, compelling models to verify the authenticity of each reasoning step. Our tripartite evaluation protocol - Causal Validation, Grounded Narrative Synthesis, and Perturbation Rejection - exposes significant vulnerabilities in state-of-the-art VLMs such as Qwen2.5-VL-72B. LogicGaze advocates for robust, trustworthy multimodal reasoning, with all resources publicly available in an anonymized repository.

📄 PDF Abstract BibTeX arXiv:2602.00292

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

Once Upon A Time In Visualization: Understanding the Use of Textual Narratives for Causality

2020-09-06 · Arjun Choudhry, Mandar Sharma, Pramod Chundury, Thomas Kapler 외

Causality visualization can help people understand temporal chains of events, such as messages sent in a distributed system, cause and effect in a historical conflict, or the interplay between political actors over time.…

Indexing and Visualization of Climate Change Narratives Using BERT and Causal Extraction

2024-08-03 · Hiroki Sakaji, Noriyasu Kaneda

In this study, we propose a methodology to extract, index, and visualize ``climate change narratives'' (stories about the connection between causal and consequential events related to climate change). We use two natural …

Articles

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives

2024-12-14 · Ji-jun Park, Soo-joon Choi

Video captioning is a critical task in the field of multimodal machine learning, aiming to generate descriptive and coherent textual narratives for video content. While large vision-language models (LVLMs) have shown sig…

DescriptiveLanguage ModelingLanguage ModellingVideo Captioning

BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal

2026-08-20 · Yigit Ekin, Enes Sanli, Aykut Erdem, Erkut Erdem 외 arxiv

Recent advances in generative video models have significantly improved visual realism in video object removal, yet evaluation protocols still focus on masked region fidelity, treating removal as local inpainting. In real…

CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding

2026-04-15 · Ishani Mondal, Yiwen Song, Mihir Parmar, Palash Goyal 외 arxiv

Long-form visual storytelling requires maintaining continuity across shots, including consistent characters, stable environments, and smooth scene transitions. While existing generative models can produce strong individu…

Visual Storytelling