paper-with-me

홈 › Papers

Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation

2024-03-21 · Leyuan Sun, Asako Kanezaki, Guillaume Caron, Yusuke Yoshiyasu

Object-goal navigation is a crucial engineering task for the community of embodied navigation; it involves navigating to an instance of a specified object category within unseen environments. Although extensive investigations have been conducted on both end-to-end and modular-based, data-driven approaches, fully enabling an agent to comprehend the environment through perceptual knowledge and perform object-goal navigation as efficiently as humans remains a significant challenge. Recently, large language models have shown potential in this task, thanks to their powerful capabilities for knowledge extraction and integration. In this study, we propose a data-driven, modular-based approach, trained on a dataset that incorporates common-sense knowledge of object-to-room relationships extracted from a large language model. We utilize the multi-channel Swin-Unet architecture to conduct multi-task learning incorporating with multimodal inputs. The results in the Habitat simulator demonstrate that our framework outperforms the baseline by an average of 10.6% in the efficiency metric, Success weighted by Path Length (SPL). The real-world demonstration shows that the proposed approach can efficiently conduct this task by traversing several rooms. For more details and real-world demonstrations, please check our project webpage (https://sunleyuan.github.io/ObjectNav).

📄 PDF Abstract BibTeX arXiv:2403.14163

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningLanguage ModelingLanguage ModellingLarge Language ModelMulti-Task LearningObject

Similar Papers 제목 키워드 기반

Semantic Layering in Room Segmentation via LLMs

2024-03-19 · Taehyeon Kim, Byung-Cheol Min

In this paper, we introduce Semantic Layering in Room Segmentation via LLMs (SeLRoS), an advanced method for semantic room segmentation by integrating Large Language Models (LLMs) with traditional 2D map-based segmentati…

Segmentation

Language and Visual Entity Relationship Graph for Agent Navigation

2020-10-19 · NeurIPS 2020 12 · Yicong Hong, Cristian Rodriguez-Opazo, Yuankai Qi, Qi Wu 외

Vision-and-Language Navigation (VLN) requires an agent to navigate in a real-world environment following natural language instructions. From both the textual and visual perspectives, we find that the relationships among …

Dynamic Time WarpingNavigateTest unseenVision and Language Navigation

CLUE: Adaptively Prioritized Contextual Cues by Leveraging a Unified Semantic Map for Effective Zero-Shot Object-Goal Navigation

2026-05-19 · Taeyun Kim, Alvin Jinsung Choi, Dasol Hong, Hyun Myung arxiv

Zero-shot object-goal navigation (ZSON) is a challenging problem in robotics that requires a comprehensive understanding of both language and visual observations. Contextual cues from rooms and objects are critical, but …

SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

2024-12-10 · Arijit Ray, Jiafei Duan, Ellis Brown, Reuben Tan 외

Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal language models (MLMs) struggle to reason ab…

Action RecognitionSpatial Reasoning

Traversability-aware Consistent Situational Graphs for Indoor Localization and Mapping

2025-10-17 · Jeewon Kim, Minho Oh, Hyun Myung arxiv

Scene graphs enhance 3D mapping capabilities in robotics by understanding the relationships between different spatial elements, such as rooms and objects. Recent research extends scene graphs to hierarchical layers, addi…

Computational Efficiency