paper-with-me

홈 › Papers

Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation

2024-06-14 · Zihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu, Shuqiang Jiang

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location in 3D environments following the natural language instruction. In this field, the agent is usually trained and evaluated in the navigation simulators, lacking effective approaches for sim-to-real transfer. The VLN agents with only a monocular camera exhibit extremely limited performance, while the mainstream VLN models trained with panoramic observation, perform better but are difficult to deploy on most monocular robots. For this case, we propose a sim-to-real transfer approach to endow the monocular robots with panoramic traversability perception and panoramic semantic understanding, thus smoothly transferring the high-performance panoramic VLN models to the common monocular robots. In this work, the semantic traversable map is proposed to predict agent-centric navigable waypoints, and the novel view representations of these navigable waypoints are predicted through the 3D feature fields. These methods broaden the limited field of view of the monocular robots and significantly improve navigation performance in the real world. Our VLN system outperforms previous SOTA monocular VLN methods in R2R-CE and RxR-CE benchmarks within the simulation environments and is also validated in real-world environments, providing a practical and high-performance solution for real-world VLN.

📄 PDF Abstract BibTeX arXiv:2406.09798

Code (1)

MrZihan/Sim2Real-VLN-3DFF 공식 구현 pytorch

Tasks

NavigateVision and Language Navigation

Similar Papers 제목 키워드 기반

Semantic-Contact Fields for Category-Level Generalizable Tactile Tool Manipulation

2026-02-14 · Kevin Yuchen Ma, Heng Zhang, Weisi Lin, Mike Zheng Shou 외 arxiv

Generalizing tool manipulation requires both semantic planning and precise physical control. Modern generalist robot policies, such as Vision-Language-Action (VLA) models, often lack the physical grounding required for c…

StyleDyRF: Zero-shot 4D Style Transfer for Dynamic Neural Radiance Fields

2024-03-13 · Hongbin Xu, Weitao Chen, Feng Xiao, Baigui Sun 외

4D style transfer aims at transferring arbitrary visual style to the synthesized novel views of a dynamic 4D scene with varying viewpoints and times. Existing efforts on 3D style transfer can effectively combine the visu…

NeRFStyle Transfer

Data-Efficient Inference of Neural Fluid Fields via SciML Foundation Model

2024-12-18 · Yuqiu Liu, Jingxuan Xu, Mauricio Soroco, Yunchao Wei 외

Recent developments in 3D vision have enabled successful progress in inferring neural fluid fields and realistic rendering of fluid dynamics. However, these methods require real-world flow captures, which demand dense vi…

g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks

2024-11-26 · CVPR 2025 1 · Zihan Wang, Gim Hee Lee

We introduce Generalizable 3D-Language Feature Fields (g3D-LF), a 3D representation model pre-trained on large-scale 3D-language dataset for embodied tasks. Our g3D-LF processes posed RGB-D images from agents to encode f…

Contrastive LearningQuestion AnsweringVision and Language Navigation

Finding 3D Scene Analogies with Multimodal Foundation Models

2025-10-27 · Junho Kim, Young Min Kim arxiv

Connecting current observations with prior experiences helps robots adapt and plan in new, unseen 3D environments. Recently, 3D scene analogies have been proposed to connect two 3D scenes, which are smooth maps that alig…