paper-with-me

홈 › Papers

ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments

2025-04-14 · Lu Yue, Dongliang Zhou, Liang Xie, Erwei Yin, Feitian Zhang

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to navigate unknown, continuous spaces based on natural language instructions. Compared to discrete settings, VLN-CE poses two core perception challenges. First, the absence of predefined observation points leads to heterogeneous visual memories and weakened global spatial correlations. Second, cumulative reconstruction errors in three-dimensional scenes introduce structural noise, impairing local feature perception. To address these challenges, this paper proposes ST-Booster, an iterative spatiotemporal booster that enhances navigation performance through multi-granularity perception and instruction-aware reasoning. ST-Booster consists of three key modules -- Hierarchical SpatioTemporal Encoding (HSTE), Multi-Granularity Aligned Fusion (MGAF), and ValueGuided Waypoint Generation (VGWG). HSTE encodes long-term global memory using topological graphs and captures shortterm local details via grid maps. MGAF aligns these dualmap representations with instructions through geometry-aware knowledge fusion. The resulting representations are iteratively refined through pretraining tasks. During reasoning, VGWG generates Guided Attention Heatmaps (GAHs) to explicitly model environment-instruction relevance and optimize waypoint selection. Extensive comparative experiments and performance analyses are conducted, demonstrating that ST-Booster outperforms existing state-of-the-art methods, particularly in complex, disturbance-prone environments.

📄 PDF Abstract BibTeX arXiv:2504.09843

Code (0)

등록된 구현이 없습니다.

Tasks

NavigateVision and Language Navigation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

FusionBooster: A Unified Image Fusion Boosting Paradigm

2023-05-10 · Chunyang Cheng, Tianyang Xu, Xiao-Jun Wu, Hui Li 외

In recent years, numerous ideas have emerged for designing a mutually reinforcing mechanism or extra stages for the image fusion task, ignoring the inevitable gaps between different vision tasks and the computational bur…

Reliable Label Correction is a Good Booster When Learning with Extremely Noisy Labels

2022-04-30 · Kai Wang, Xiangyu Peng, Shuo Yang, Jianfei Yang 외

Learning with noisy labels has aroused much research interest since data annotations, especially for large-scale datasets, may be inevitably imperfect. Recent approaches resort to a semi-supervised learning problem by di…

Learning with noisy labels

EG-Booster: Explanation-Guided Booster of ML Evasion Attacks

2021-08-31 · Abderrahmen Amich, Birhanu Eshete

The widespread usage of machine learning (ML) in a myriad of domains has raised questions about its trustworthiness in security-critical environments. Part of the quest for trustworthy ML is robustness evaluation of ML m…

image-classificationImage Classification

TrajBooster: Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning

2025-09-15 · Jiacheng Liu, Pengxiang Ding, Qihang Zhou, Yuxuan Wu 외 arxiv

Recent Vision-Language-Action models show potential to generalize across embodiments but struggle to quickly align with a new robot's action space when high-quality demonstrations are scarce, especially for bipedal human…

UADB: Unsupervised Anomaly Detection Booster

2023-06-03 · Hangting Ye, Zhining Liu, Xinyi Shen, Wei Cao 외

Unsupervised Anomaly Detection (UAD) is a key data mining problem owing to its wide real-world applications. Due to the complete absence of supervision signals, UAD methods rely on implicit assumptions about anomalous pa…

Anomaly DetectionUnsupervised Anomaly Detection