paper-with-me

홈 › Papers

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

2026-06-02 · Yifei Li, Pengyiang Liu, Yuhang Zang, Zhongyue Shi, Qi Fu, Hongye Hao, Jiwen Lu arxiv

Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view. Existing benchmarks either evaluate offline over full videos or target events rather than spatial structure. We introduce OVO-S-Bench, a fully human-annotated benchmark for streaming spatial intelligence, comprising 1,680 questions over 348 source videos. Annotation involves 12 trained annotators (each also serving as a blind cross-reviewer) across roughly 804 person-hours of multi-round quality assurance. Each question carries a query timestamp and an evidence interval, and at evaluation, the model sees only the prefix preceding the query. Questions span four levels of increasing abstraction: instantaneous egocentric perception, spatiotemporal context tracking, generative spatial reasoning, and allocentric spatial mapping. Across 38 proprietary and open-source MLLMs, Gemini-3.1-Pro trails human experts evaluated under the same prefix-access protocol by 33 points (59.2 vs. 92.2), with allocentric spatial mapping as the dominant bottleneck. Notably, streaming and spatially fine-tuned MLLMs underperform their own backbones. We further find that chain-of-thought reasoning amplifies spatial errors when ungrounded in the stream. By exposing these limitations, OVO-S-Bench establishes a demanding testbed for next-generation streaming spatial MLLMs.

📄 PDF Abstract BibTeX arXiv:2606.03890

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

2026-06-13 · Yun Wang, Junbin Xiao, Han Lyu, Yifan Wang 외 arxiv

We introduce UCS-Bench, a dataset spanning 170+ hours of egocentric visual observations with 8.1K+ timestamped questions for diagnosing User-Centric Continual Spatial intelligence in egocentric video streams. UCS-Bench t…

Spatial Reasoning

FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding

2026-03-02 · Yiweng Xie, Bo He, Junke Wang, Xiangyu Zheng 외 arxiv

This paper presents FluxMem, a training-free framework for efficient streaming video understanding. FluxMem adaptively compresses redundant visual memory through a hierarchical, two-stage design: (1) a Temporal Adjacency…

Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training

2026-03-12 · Fangfu Liu, Diankun Wu, Jiawei Chi, Yimo Cai 외 arxiv

Humans perceive and understand real-world spaces through a stream of visual observations. Therefore, the ability to streamingly maintain and update spatial evidence from potentially unbounded video streams is essential f…

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

2025-11-26 · Peiran Xu, Sudong Wang, Yao Zhu, Jianing Li 외 arxiv

Spatial cognition is fundamental to real-world multimodal intelligence, allowing models to effectively interact with the physical environment. While multimodal large language models (MLLMs) have made significant strides,…

Spatial ReasoningCausal Inference

See, Remember, Explore: A Benchmark and Baselines for Streaming Spatial Reasoning

2026-03-25 · Yuxi Wei, Wei Huang, Qirui Chen, Lu Hou 외 arxiv

Spatial understanding is fundamental for embodied agents, yet most spatial VLMs and benchmarks remain offline-evaluating post-hoc QA over pre-recorded inputs and overlooking two crucial deployment-critical requirements: …

Question AnsweringSpatial Reasoning