paper-with-me

홈 › Papers

Chronological Blindness: Benchmarking Temporal Reasoning in Vision-Language Models with CHRONOSIGHT

2026-06-15 · Parthaw Goswami, Jaynto Goswami Deep arxiv

Human perception of visual scenes is inherently temporal. We instinctively recognise whether a fruit is ripening or rotting, whether construction is progressing or being demolished, and approximately how much time separates two photographs of the same subject. Whether large vision-language models (VLMs) share this competence remains an open and practically important question. We introduce CHRONOSIGHT, a rigorously controlled benchmark evaluating five dimensions of visual temporal reasoning: CHRONORANK (chronological ordering of image sequences), CHRONOLOCATE (ordinal stage localisation from a single image), CHRONODELTA (estimation of time elapsed between two images on a logarithmic scale), CHRONOREVERSE (detection of temporally reversed sequences), and CHRONOODD (identification of a temporal outlier within a set). The benchmark comprises 1{,}000 items across eight process families (biological growth, food transformation, physical weathering, construction, environmental change, human ageing, astronomical phenomena, and urban dynamics) spanning timescales from minutes to millennia. We evaluate eight open-source VLMs (500 M to 19 B parameters) under two prompting regimes and collect human performance baselines. Human performance averages 0.89 across tasks; the best open model (Qwen2.5-VL-7B) reaches 0.40 under direct prompting, a gap we term chronological blindness. Lightweight LoRA fine-tuning on 151 examples raises CHRONODELTA accuracy from near-zero to 0.43, transferring zero-shot to related tasks (CHRONOODD: 0.37; CHRONOREVERSE: 0.64)suggesting the bottleneck is partly instruction following rather than visual perception. Benchmark, code, and predictions will be released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2606.16334

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models

2026-06-04 · Haoyu Zhou, Qing Qing, Caichong Li, Qixin Zhang 외 arxiv

Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capacity for chronological reasoning remains under-explored. In this paper, …

ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?

2025-09-23 · Zijian Ling, Han Zhang, Yazhuo Zhou, Jiahao Cui arxiv

This paper presents ColorBlindnessEval, a novel benchmark designed to evaluate the robustness of Vision-Language Models (VLMs) in visually adversarial scenarios inspired by the Ishihara color blindness test. Our dataset …

MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models

2026-05-19 · Yuanqing Cai, Ziyi Huang, Minhao Liu, Lixin Duan 외 arxiv

Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of \emph{inattentional blindness} in human cognition, we investigate whether LLMs, trained on human-prefer…

From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning

2026-04-12 · Xiaoda Yang, Yuxiang Liu, Shenzhou Gao, Can Wang 외 arxiv

Modern vision-language models achieve strong performance in static perception, but remain limited in the complex spatiotemporal reasoning required for embodied, egocentric tasks. A major source of failure is their relian…

Logical Reasoning

Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents

2025-12-23 · Yiming Du, Baojun Wang, Yifan Xiang, Zhaowei Wang 외 arxiv

Temporal reasoning over long, multi-session dialogues is a critical capability for conversational agents. However, existing works and our pilot study have shown that as dialogue histories grow in length and accumulate no…

Reinforcement Learning