paper-with-me

홈 › Papers

Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models

2026-05-09 · Tri Cao, Khoi Le, Thong Nguyen, Cong-Duy Nguyen, Quynh Vo, Anh Tuan Luu, Chunyan Miao, See-Kiong Ng, Shuicheng Yan, Bryan Hooi arxiv

While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure in spatio-temporal monitoring, the ability to persistently track object identities, states, and relations over time. Existing benchmarks obscure this deficit by relying on single final-answer evaluations for queries that can often be resolved via local visual cues or statistical priors. To rigorously diagnose this, we introduce STEMO-Bench (Spatio-TEmporal MOnitoring), a benchmark of human-verified object-centric facts that evaluates intermediate reasoning by decomposing queries into sub-questions, distinguishing genuine temporal understanding from coincidental correctness. To address failure modes exposed by STEMO, we propose STEMO-Track, a novel object-centric framework that explicitly constructs and reasons over structured object trajectories via chunk-wise state extraction and temporal aggregation. Extensive experiments demonstrate that our object-centric framework significantly reduces hallucinated answers and improves spatio-temporal reasoning consistency over state-of-the-art MLLMs.

📄 PDF Abstract BibTeX arXiv:2605.08974

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SiamMo: Siamese Motion-Centric 3D Object Tracking

2024-08-03 · Yuxiang Yang, Yingqi Deng, Jing Zhang, Hongjie Gu 외

Current 3D single object tracking methods primarily rely on the Siamese matching-based paradigm, which struggles with textureless and incomplete LiDAR point clouds. Conversely, the motion-centric paradigm avoids appearan…

3D Object Tracking3D Single Object TrackingMotion EstimationObject+1

EgoTracks: A Long-term Egocentric Visual Object Tracking Dataset

2023-01-09 · NeurIPS 2023 11

Visual object tracking is a key component to many egocentric vision problems. However, the full spectrum of challenges of egocentric tracking faced by an embodied AI is underrepresented in many existing datasets; these t…

ObjectObject TrackingVisual Object Tracking

Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

2025-04-07 · Yunlong Tang, Jing Bi, Chao Huang, Susan Liang 외

We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects through time. CAT-V integrates three ke…

Boundary DetectionObjectSemantic SegmentationVideo Captioning

Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification

2026-06-27 · Tianshu Zhang, Yan Wang, Ji Qi, Lijie Wen arxiv

Spatio-temporal grounding in long videos requires precise temporal localization and robust object tracking conditioned on natural-language queries. While recent vision-language models (VLMs) show strong reasoning ability…

Reinforcement LearningObject Tracking

SpOT: Spatiotemporal Modeling for 3D Object Tracking

2022-07-12 · Colton Stearns, Davis Rempe, Jie Li, Rares Ambrus 외

3D multi-object tracking aims to uniquely and consistently identify all mobile entities through time. Despite the rich spatiotemporal information available in this setting, current 3D tracking methods primarily rely on a…

3D Multi-Object Tracking3D Object TrackingMulti-Object TrackingObject+1