paper-with-me

홈 › Papers

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models

2026-05-11 · Junzhe Chen, Siyuan Meng, Yuxi Chen, Man Zhao, Wenyao Gui, Xiaojie Guo arxiv

Video large language models (Video-LLMs) have made strong progress in general video understanding, but their ability to maintain temporal object consistency remains underexplored. Existing benchmarks often emphasize event recognition, action understanding, or coarse temporal reasoning, while rarely testing whether models can preserve the identity, state, and continuity of the same object across occlusion, disappearance, reappearance, state transitions, and cross-object interactions. We introduce TOC-Bench, a diagnostic benchmark for evaluating temporal object consistency in Video-LLMs. TOC-Bench is object-track grounded: each queried subject is linked to a per-frame trajectory and a structured temporal event timeline. To ensure that questions require temporally ordered visual evidence rather than language priors, single-frame shortcuts, or unordered frame cues, we design a three-layer temporal-necessity filtering protocol, which removes 60.7% of candidate QA pairs and retains 17,900 temporally dependent items across 10 diagnostic dimensions. From this pool, we construct a human-verified benchmark with 2,323 high-quality QA pairs over 1,951 videos. Experiments on representative Video-LLMs show that temporal object consistency remains a major unsolved challenge, with notable weaknesses in event counting, event ordering, identity-sensitive reasoning, and hallucination-aware verification, even when models perform well on general video understanding benchmarks. These results suggest that object-centric temporal coherence is a key bottleneck for current Video-LLMs, and that TOC-Bench provides a focused platform for diagnosing and improving object-aware temporal reasoning. The resource is available at https://github.com/cjzcjz666/toc_bench.git.

📄 PDF Abstract BibTeX arXiv:2605.09904

Code (0)

등록된 구현이 없습니다.

Tasks

Action Understanding

Similar Papers 제목 키워드 기반

Temporally Consistent Referring Video Object Segmentation with Hybrid Memory

2024-03-28 · Bo Miao, Mohammed Bennamoun, Yongsheng Gao, Mubarak Shah 외

Referring Video Object Segmentation (R-VOS) methods face challenges in maintaining consistent object segmentation due to temporal context variability and the presence of other visually similar objects. We propose an end-…

HTRObjectReferring Expression SegmentationReferring Video Object Segmentation+5

IDENTIFYING CONCEALED OBJECTS FROM VIDEOS

2021-09-29 · Xuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong 외

Concealed objects are often hard to identify from still images, as often camouflaged objects exhibit patterns seamless to the background. In this work, we propose a novel video concealed object detection (VCOD) framework…

object-detectionObject Detection

FiVE: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models

2025-03-17 · Minghan Li, Chenxi Xie, Yichen Wu, Lei Zhang 외

Numerous text-to-video (T2V) editing methods have emerged recently, but the lack of a standardized benchmark for fair evaluation has led to inconsistent claims and an inability to assess model sensitivity to hyperparamet…

SensitivityVideo EditingVideo Similarity

Learn Temporal Consistency For Robust Satellite Video Detector

2026-06-13 · Weilong Guo, Shengyang Li, Yanfeng Gu arxiv

Satellite video object detection (SVOD) for oriented and fine-grained objects plays an important role in satellite applications. Most existing SVOD methods only focus on one or a few coarse-grained categories of moving o…

Representation LearningVideo Object Detection

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales

2026-06-20 · Xinlong Chen, Jiafu Tang, Yue Ding, Yizhuo Jia 외 arxiv

Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks. However, few existing benchmarks can objectively and comprehensively evaluate t…

Video Captioning