paper-with-me

홈 › Papers

UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation

2024-11-25 · Guangzhao Dai, Jian Zhao, Yuantao Chen, Yusen Qin, Hao Zhao, GuoSen Xie, Yazhou Yao, Xiangbo Shu, Xuelong Li

Vision-and-Language Navigation (VLN), where an agent follows instructions to reach a target destination, has recently seen significant advancements. In contrast to navigation in discrete environments with predefined trajectories, VLN in Continuous Environments (VLN-CE) presents greater challenges, as the agent is free to navigate any unobstructed location and is more vulnerable to visual occlusions or blind spots. Recent approaches have attempted to address this by imagining future environments, either through predicted future visual images or semantic features, rather than relying solely on current observations. However, these RGB-based and feature-based methods lack intuitive appearance-level information or high-level semantic complexity crucial for effective navigation. To overcome these limitations, we introduce a novel, generalizable 3DGS-based pre-training paradigm, called UnitedVLN, which enables agents to better explore future environments by unitedly rendering high-fidelity 360 visual images and semantic features. UnitedVLN employs two key schemes: search-then-query sampling and separate-then-united rendering, which facilitate efficient exploitation of neural primitives, helping to integrate both appearance and semantic information for more robust navigation. Extensive experiments demonstrate that UnitedVLN outperforms state-of-the-art methods on existing VLN-CE benchmarks.

📄 PDF Abstract BibTeX arXiv:2411.16053

Code (0)

등록된 구현이 없습니다.

Tasks

3DGSNavigateVision and Language NavigationVision-Language Navigation

Similar Papers 제목 키워드 기반

SatSurfGS: Generalizable 2D Gaussian Splatting for Sparse-View Satellite Surface Reconstruction

2026-05-08 · Min Chen, Wei Guo, Bin Wang, Wen Li 외 arxiv

Sparse-view satellite image surface reconstruction remains highly challenging, fundamentally because the reliability of multi-view matching under satellite imaging conditions is strongly spatially heterogeneous. Affected…

SmileSplat: Generalizable Gaussian Splats for Unconstrained Sparse Images

2024-11-27 · Yanyan Li, Yixin Fang, Federico Tombari, Gim Hee Lee

Sparse Multi-view Images can be Learned to predict explicit radiance fields via Generalizable Gaussian Splatting approaches, which can achieve wider application prospects in real-life when ground-truth camera parameters …

DecoderNovel View Synthesis

GPS-Gaussian+: Generalizable Pixel-wise 3D Gaussian Splatting for Real-Time Human-Scene Rendering from Sparse Views

2024-11-18 · Boyao Zhou, Shunyuan Zheng, Hanzhang Tu, Ruizhi Shao 외

Differentiable rendering techniques have recently shown promising results for free-viewpoint video synthesis of characters. However, such methods, either Gaussian Splatting or neural implicit rendering, typically necessi…

Depth EstimationNovel View Synthesis

HFGaussian: Learning Generalizable Gaussian Human with Integrated Human Features

2024-11-05 · Arnab Dey, Cheng-You Lu, Andrew I. Comport, Srinath Sridhar 외

Recent advancements in radiance field rendering show promising results in 3D scene representation, where Gaussian splatting-based techniques emerge as state-of-the-art due to their quality and efficiency. Gaussian splatt…

Feature SplattingPose Estimation

Reinforcement Learning with Generalizable Gaussian Splatting

2024-03-18 · Jiaxu Wang, Qiang Zhang, Jingkai Sun, Jiahang Cao 외

An excellent representation is crucial for reinforcement learning (RL) performance, especially in vision-based reinforcement learning tasks. The quality of the environment representation directly influences the achieveme…

3DGSreinforcement-learningReinforcement LearningReinforcement Learning (RL)