paper-with-me

홈 › Papers

SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language Navigation

2021-10-27 · NeurIPS 2021 12 · Abhinav Moudgil, Arjun Majumdar, Harsh Agrawal, Stefan Lee, Dhruv Batra

Natural language instructions for visual navigation often use scene descriptions (e.g., "bedroom") and object references (e.g., "green chairs") to provide a breadcrumb trail to a goal location. This work presents a transformer-based vision-and-language navigation (VLN) agent that uses two different visual encoders -- a scene classification network and an object detector -- which produce features that match these two distinct types of visual cues. In our method, scene features contribute high-level contextual information that supports object-level processing. With this design, our model is able to use vision-and-language pretraining (i.e., learning the alignment between images and text from large-scale web data) to substantially improve performance on the Room-to-Room (R2R) and Room-Across-Room (RxR) benchmarks. Specifically, our approach leads to improvements of 1.8% absolute in SPL on R2R and 3.7% absolute in SR on RxR. Our analysis reveals even larger gains for navigation instructions that contain six or more object references, which further suggests that our approach is better able to use object features and align them to references in the instructions.

📄 PDF Abstract BibTeX arXiv:2110.14143

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectScene ClassificationVision and Language NavigationVisual Navigation

Similar Papers 제목 키워드 기반

Revisiting Attention for Multivariate Time Series Forecasting

2024-07-18 · Haixiang Wu

Current Transformer methods for Multivariate Time-Series Forecasting (MTSF) are all based on the conventional attention mechanism. They involve sequence embedding and performing a linear projection of Q, K, and V, and th…

Multivariate Time Series ForecastingTime SeriesTime Series Forecasting

Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies

2024-05-24 · Jianing Qian, Anastasios Panagopoulos, Dinesh Jayaraman

Generic re-usable pre-trained image representation encoders have become a standard component of methods for many computer vision tasks. As visual representations for robots however, their utility has been limited, leadin…

Neural Attention Field: Emerging Point Relevance in 3D Scenes for One-Shot Dexterous Grasping

2024-10-30 · Qianxu Wang, Congyue Deng, Tyler Ga Wei Lum, Yuanpei Chen 외

One-shot transfer of dexterous grasps to novel scenes with object and context variations has been a challenging problem. While distilled feature fields from large vision models have enabled semantic correspondences acros…

Decoder

Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications

2025-03-25 · Ben Rahman

Semantic segmentation has made significant strides in pixel-level image understanding, yet it remains limited in capturing contextual and semantic relationships between objects. Current models, such as CNN and Transforme…

Autonomous DrivingSemantic Segmentation

Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios

2025-10-30 · Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi arxiv

In real-world environments, AI systems often face unfamiliar scenarios without labeled data, creating a major challenge for conventional scene understanding models. The inability to generalize across unseen contexts limi…

Zero-shot GeneralizationScene Understanding