paper-with-me

Papers

Cog3DMap: Multi-View Vision-Language Reasoning with 3D Cognitive Maps

2026-03-24 · Chanyoung Gwak, Yoonwoo Jeong, Byungwoo Jeon, Hyunseok Lee, Jinwoo Shin, Minsu Cho arxiv

Precise spatial understanding from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs), as their visual representations are predominantly semantic and lack explicit geometric grounding. While existing approaches augment visual tokens with geometric cues from visual geometry models, their MLLM is still required to implicitly infer the underlying 3D structure of the scene from these augmented tokens, limiting its spatial reasoning capability. To address this issue, we introduce Cog3DMap, a framework that recurrently constructs an explicit 3D memory from multi-view images, where each token is grounded in 3D space and possesses both semantic and geometric information. By feeding these tokens into the MLLM, our framework enables direct reasoning over a spatially structured 3D map, achieving state-of-the-art performance on various spatial reasoning benchmarks. Code will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2603.23023

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

2025-06-30 · Zhaochen Su, Peng Xia, Hangyu Guo, Zhenhua Liu 외

Recent progress in multimodal reasoning has been significantly advanced by textual Chain-of-Thought (CoT), a paradigm where models conduct reasoning within language. This text-centric approach, however, treats vision as …

Multimodal Reasoning

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap

2026-04-15 · Hanxuan Chen, Jie Zheng, Siqi Yang, Tianle Zeng 외 arxiv

Vision-and-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) represents a pivotal challenge in embodied artificial intelligence, focused on enabling UAVs to interpret high-level human commands and execute long-h…

WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models

2026-04-09 · Hongjin Chen, Shangyun Jiang, Tonghua Su, Chen Gao 외 arxiv

Vision-language models (VLMs) and generative world models are opening new opportunities for embodied navigation. VLMs are increasingly used as direct planners or trajectory predictors, while world models support look-ahe…

Vision-Language NavigationTrajectory Prediction

RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning

2025-10-02 · Sicheng Feng, Kaiwen Tuo, Song Wang, Lingdong Kong 외 arxiv

Fine-grained visual reasoning remains a core challenge for multimodal large language models (MLLMs). The recently introduced ReasonMap highlights this gap by showing that even advanced MLLMs struggle with spatial reasoni…

Visual Question AnsweringReinforcement LearningSpatial ReasoningVisual Reasoning

Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control

2026-06-27 · Runji Cai, Toshihiko Yamasaki, Ling Xiao arxiv

Social robot navigation (SRN) requires more than geometric path planning; it demands understanding human intentions, social norms, and contextual cues to generate socially compliant behaviors. Although classical navigati…

Collision AvoidanceRobot Navigation