paper-with-me

홈 › Papers

Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability

2025-10-09 · Chengzhi Li, Heyan Huang, Ping Jian, Zhen Yang, Yaning Tian, Zhongbin Guo arxiv

Large language models (LLMs) often generate self-contradictory outputs, which severely impacts their reliability and hinders their adoption in practical applications. In video-language models (Video-LLMs), this phenomenon recently draws the attention of researchers. Specifically, these models fail to provide logically consistent responses to rephrased questions based on their grounding outputs. However, the underlying causes of this phenomenon remain underexplored. In this work, we adopt an interpretability-driven approach to analyze, statistically summarize, and intervention the potential factors of the phenomenon. We find that one of the primary reasons for the inconsistency in responses lies in the inability of cross-modal attention heads to effectively distinguish video tokens across different timestamps. To address this, we propose an attention enhancement method called Temporally Conditioned Attention Sharpening (TCAS), which constructs an enhancement objective based on attention distinctions to enhance the model's temporal resolution capability, thereby improving its temporal understanding logic consistency. Experimental results demonstrate that our method significantly enhances the temporal logic consistency of Video-LLMs. Further analyses reveal that our method indeed improves the temporal discriminability of attention heads, validating our conclusions. Additionally, our method even achieves performance improvements in general video temporal grounding tasks, suggesting that temporal logic consistency is an important factor in temporal understanding.

📄 PDF Abstract BibTeX arXiv:2510.08138

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models

2026-05-11 · Junzhe Chen, Siyuan Meng, Yuxi Chen, Man Zhao 외 arxiv

Video large language models (Video-LLMs) have made strong progress in general video understanding, but their ability to maintain temporal object consistency remains underexplored. Existing benchmarks often emphasize even…

Action Understanding

TimeLogic: A Temporal Logic Benchmark for Video QA

2025-01-13 · Sirnam Swetha, Hilde Kuehne, Mubarak Shah

Temporal logical understanding, a core facet of human cognition, plays a pivotal role in capturing complex sequential events and their temporal relationships within videos. This capability is particularly crucial in task…

2kAction SegmentationLogical ReasoningQuestion Answering+2

Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models

2025-11-28 · Muhammad Maaz, Hanoona Rasheed, Fahad Shahbaz Khan, Salman Khan arxiv

Reasoning over dynamic visual content remains a central challenge for multimodal large language models. Recent thinking models generate explicit reasoning traces for interpretability; however, their reasoning often appea…

Reinforcement Learning

Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization

2025-03-22 · Zhuo Tao, Liang Li, Qi Chen, Yunbin Tu 외

Natural language video localization (NLVL) is a crucial task in video understanding that aims to localize the target moment in videos specified by a given language description. Recently, a point-supervised paradigm has b…

Saliency DetectionSentenceVideo Understanding

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

2025-10-30 · Minjoon Jung, Junbin Xiao, Junghyun Kim, Byoung-Tak Zhang 외 arxiv

Do Video-LLMs have consistent temporal understanding when videos capture the same event from different viewpoints? To study this question, we introduce EgoExo-Con(sistency), a benchmark of synchronized egocentric and exo…

Reinforcement Learning