paper-with-me

홈 › Papers

Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation

2025-10-01 · Yunbo Xu, Xuesong Zhang, Jia Li, Zhenzhen Hu, Richang Hong arxiv

Following language instructions, vision-language navigation (VLN) agents are tasked with navigating unseen environments. While augmenting multifaceted visual representations has propelled advancements in VLN, the significance of foreground and background in visual observations remains underexplored. Intuitively, foreground regions provide semantic cues, whereas the background encompasses spatial connectivity information. Inspired on this insight, we propose a Consensus-driven Online Feature Augmentation strategy (COFA) with alternative foreground and background features to facilitate the navigable generalization. Specifically, we first leverage semantically-enhanced landmark identification to disentangle foreground and background as candidate augmented features. Subsequently, a consensus-driven online augmentation strategy encourages the agent to consolidate two-stage voting results on feature preferences according to diverse instructions and navigational locations. Experiments on REVERIE and R2R demonstrate that our online foreground-background augmentation boosts the generalization of baseline and attains state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2510.00604

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language Navigation

Similar Papers 제목 키워드 기반

RARA: Zero-shot Sim2Real Visual Navigation with Following Foreground Cues

2022-01-08 · Klaas Kelchtermans, Tinne Tuytelaars

The gap between simulation and the real-world restrains many machine learning breakthroughs in computer vision and reinforcement learning from being applicable in the real world. In this work, we tackle this gap for the …

TripletVisual Navigation

Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning

2025-09-18 · Wenda Qin, Andrea Burns, Bryan A. Plummer, Margrit Betke arxiv

Large models achieve strong performance on Vision-and-Language Navigation (VLN) tasks, but are costly to run in resource-limited environments. Token pruning offers appealing tradeoffs for efficiency with minimal performa…

Noise or Signal: The Role of Image Backgrounds in Object Recognition

2020-06-17 · ICLR 2021 1 · Kai Xiao, Logan Engstrom, Andrew Ilyas, Aleksander Madry

We assess the tendency of state-of-the-art object recognition models to depend on signals from image backgrounds. We create a toolkit for disentangling foreground and background signal on ImageNet images, and find that (…

BIG-bench Machine LearningObject Recognition

Unsupervised Foreground Extraction via Deep Region Competition

2021-10-29 · NeurIPS 2021 12 · Peiyu Yu, Sirui Xie, Xiaojian Ma, Yixin Zhu 외

We present Deep Region Competition (DRC), an algorithm designed to extract foreground objects from images in a fully unsupervised manner. Foreground extraction can be viewed as a special case of generic image segmentatio…

Image SegmentationInductive BiasMixture-of-ExpertsSemantic Segmentation

Volumetric Disentanglement for 3D Scene Manipulation

2022-06-06 · Sagie Benaim, Frederik Warburg, Peter Ebert Christensen, Serge Belongie

Recently, advances in differential volumetric rendering enabled significant breakthroughs in the photo-realistic and fine-detailed reconstruction of complex 3D scenes, which is key for many virtual reality applications. …

DisentanglementObject