paper-with-me

홈 › Papers

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation

2026-05-17 · Jingzhi Huang, Junkai Huang, Wenxuan Song, Haoyang Yang, Hailong Huang, Haoang Li, Yi Wang arxiv

Vision-Language Navigation (VLN) approaches have currently followed two primary paradigms: the end-to-end Vision-Language Model (VLM) policy fine-tuned on navigation trajectories to directly predict actions, and the zero-shot modular pipeline integrating pre-trained Multimodal Large Language Model (MLLM) for training-free generalization to unseen environments. However, end-to-end methods struggle with long-horizon navigation and lack dynamic reasoning, whereas zero-shot methods are constrained by limited spatial grounding for reliable planning and also require substantial reasoning time. To bridge this gap, we introduce SEDualVLN, a spatially-enhanced dual-system VLN framework. System 1 is a VLM model enhanced with both global and local spatial awareness, used for action generation. System 2 integrates a general MLLM with a mapping module, wherein the MLLM plans waypoints by leveraging top-down views of the real-time 3D map alongside streams of rendered path images. Both systems leverage different forms of spatial enhancement to cultivate the agent's sense of direction in VLN tasks. Ultimately, they cooperate to complete the navigation task through a fast-slow coordinated approach. SEDualVLN achieves state-of-the-art performance on VLN-CE benchmarks, and further ablation studies demonstrate the effectiveness of each system and module.

📄 PDF Abstract BibTeX arXiv:2605.17249

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language Navigation

Similar Papers 제목 키워드 기반

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

2026-05-05 · Lin Song, Wenbo Li, Guoqing Ma, Wei Tang 외 arxiv

We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatially enhanced Multimodal Large Language M…

Text-to-Image GenerationImage Editing

Multipath-Enhanced Device-Free Localization in Wideband Wireless Networks

2020-10-09 · Martin Schmidhammer, Christian Gentner, Stephan Sand, Uwe-Carsten Fiebig

State-of-the-art device-free localization systems infer presence and location of users based on received signal strength measurements of line-of-sight links in wireless networks. In this letter, we propose to enhance dev…

A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision

2025-05-16 · Alexey Magay, Dhurba Tripathi, Yu Hao, Yi Fang

People with blindness and low vision (pBLV) face significant challenges, struggling to navigate environments and locate objects due to limited visual cues. Spatial reasoning is crucial for these individuals, as it enable…

Large Language ModelNavigateObject RecognitionSpatial Reasoning

RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence

2025-12-31 · Chengkai Hou, Kun Wu, Jiaming Liu, Zhengping Che 외 arxiv

While data-driven imitation learning has revolutionized robotic manipulation, current approaches remain constrained by the scarcity of large-scale, diverse real-world demonstrations. Consequently, the ability of existing…

Reinforcement Learning

Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation

2024-12-08 · CVPR 2025 1 · Hyeonho Jeong, Chun-Hao Paul Huang, Jong Chul Ye, Niloy Mitra 외

While recent foundational video generators produce visually rich output, they still struggle with appearance drift, where objects gradually degrade or change inconsistently across frames, breaking visual coherence. We hy…

Point TrackingVideo Generation