paper-with-me

홈 › Papers

Spatial-Conditioned Reasoning in Long-Egocentric Videos

2026-01-26 · James Tribble, Hao Wang, Si-En Hong, Chaoyi Zhou, Ashish Bastola, Siyu Huang, Abolfazl Razi arxiv

Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-language models perform well on image and short-video reasoning, their spatial reasoning capability in long egocentric sequences remains limited. In this work, we study how explicit spatial signals influence VLM-based video understanding without modifying model architectures or inference procedures. We introduce Sanpo-D, a fine-grained re-annotation of the Google Sanpo dataset, and benchmark multiple VLMs on navigation-oriented spatial queries. To examine input-level inductive bias, we further fuse depth maps with RGB frames and evaluate their impact on spatial reasoning. Our results reveal a trade-off between general-purpose accuracy and spatial specialization, showing that depth-aware and spatially grounded representations can improve performance on safety-critical tasks such as pedestrian and obstruction detection.

📄 PDF Abstract BibTeX arXiv:2601.18100

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningVisual Navigation

Similar Papers 제목 키워드 기반

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

2026-05-18 · Kunyu Peng, Zhikun Zhou, Kailun Yang, Di Wen 외 arxiv

Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoints remains largely unexplored. We study …

Spatial Reasoning

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs

2026-06-24 · Agnese Taluzzi, Riccardo Santambrogio, Simone Mentasti, Chiara Plizzari 외 arxiv

Existing multi-modal large language models (MLLMs) face significant challenges in processing long video sequences due to strict input token limitations. As a result, current video understanding approaches, especially in …

Video Question Answering

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning

2026-03-24 · Jiacheng Hua, Yishu Yin, Yuhang Wu, Tai Wang 외 arxiv

Existing Multimodal Large Language Models (MLLMs) struggle with 3D spatial reasoning, as they fail to construct structured abstractions of the 3D environment depicted in video inputs. To bridge this gap, drawing inspirat…

Question AnsweringSpatial Reasoning

EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos

2026-05-18 · Ruiping Liu, Junwei Zheng, Yufan Chen, Di Wen 외 arxiv

Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial-temporal reasoning. Inspired by human recall from both field and observer perspectives, we introduce EgoExoM…

Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

2026-06-13 · Yun Wang, Junbin Xiao, Han Lyu, Yifan Wang 외 arxiv

We introduce UCS-Bench, a dataset spanning 170+ hours of egocentric visual observations with 8.1K+ timestamped questions for diagnosing User-Centric Continual Spatial intelligence in egocentric video streams. UCS-Bench t…

Spatial Reasoning