paper-with-me

홈 › Papers

StateTrace: An Object-Centric Framework for Hidden-State Spatiotemporal Reasoning in Long Videos

2026-08-19 · Yu Han, Wenhao Li, Yichao Cao, Hongyan Xu, Shuo Yang, Shan You, Xiu Su arxiv

Existing VLMs have achieved strong performance in video understanding, yet they struggle with long-video spatiotemporal reasoning when target objects become invisible, often mistaking "invisible" for "unknown". We define this challenge as hidden-state spatiotemporal reasoning: inferring object states during prolonged invisible intervals from context interactions. To address this, we propose StateTrace, a novel object-centric framework that endows VideoLLMs with an explicit mechanism for hidden state reasoning in long videos. StateTrace builds a reusable spatiotemporal state memory that organizes object trajectories, inter-object relations, and state-transition events into a structured reasoning substrate. At inference time, it retrieves question-relevant state-evolution trajectories and converts them into compact reasoning cues, enabling the model to explicitly reason about why an object disappears, how its state evolves while invisible, and whether that state should persist at query time. We further build HSR-Bench, a diagnostic benchmark for hidden-state reasoning, containing 1,427 video-QA samples from 1,384 unique videos. Extensive experiments across multiple VideoLLMs show that StateTrace consistently improves performance on both public benchmarks and HSR-Bench (e.g., improving VideoLLaMA3 from 39.6 to 64.2 on HSR-Bench).

📄 PDF Abstract BibTeX arXiv:2608.18532

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FlowMo-WM: A World Model with Object Momentum and Hidden Ambient Drift

2026-06-11 · Yitao Jiang, Luyang Zhao, Muhao Chen, Devin Balkcom arxiv

World models in robot learning predict future states from visual observations and actions, enabling agents to reason about the consequences of their controls. However, many action-conditioned models are evaluated in sett…

Toward a Metrology for Artificial Intelligence: Hidden-Rule Environments and Reinforcement Learning

2025-09-07 · Christo Mathew, Wentian Wang, Jacob Feldman, Lazaros K. Gallos 외 arxiv

We investigate reinforcement learning in the Game Of Hidden Rules (GOHR) environment, a complex puzzle in which an agent must infer and execute hidden rules to clear a 6$\times$6 board by placing game pieces into buckets…

Reinforcement Learning

Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning

2026-05-13 · Qinchuan Cheng, Zhantao Gong, Pengzhan Sun, Angela Yao 외 arxiv

Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmarks only partially test this ability. Egoc…

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

2025-11-14 · Nhat Chung, Taisei Hanyu, Toan Nguyen, Huy Le 외 arxiv

As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interacti…

Object Tracking

Recurrent Attention Models with Object-centric Capsule Representation for Multi-object Recognition

2021-10-11 · Hossein Adeli, Seoyoung Ahn, Gregory Zelinsky

The visual system processes a scene using a sequence of selective glimpses, each driven by spatial and object-based attention. These glimpses reflect what is relevant to the ongoing task and are selected through recurren…

DecoderObjectObject Recognition