The RobotSlang Benchmark: Dialog-guided Robot Localization and Navigation
Autonomous robot systems for applications from search and rescue to assistive guidance should be able to engage in natural language dialog with people. To study such cooperative communication, we introduce Robot Simultaneous Localization and Mapping with Natural Language (RobotSlang), a benchmark of 169 natural language dialogs between a human Driver controlling a robot and a human Commander providing guidance towards navigation goals. In each trial, the pair first cooperates to localize the robot on a global map visible to the Commander, then the Driver follows Commander instructions to move the robot to a sequence of target objects. We introduce a Localization from Dialog History (LDH) and a Navigation from Dialog History (NDH) task where a learned agent is given dialog and visual observations from the robot platform as input and must localize in the global map or navigate towards the next target object, respectively. RobotSlang is comprised of nearly 5k utterances and over 1k minutes of robot camera and control streams. We present an initial model for the NDH task, and show that an agent trained in simulation can follow the RobotSlang dialog-based navigation instructions for controlling a physical robot platform. Code and data are available at https://umrobotslang.github.io/.
Code (0)
등록된 구현이 없습니다.
Tasks
NavigateSimultaneous Localization and MappingSimilar Papers 제목 키워드 기반
GeomGS: LiDAR-Guided Geometry-Aware Gaussian Splatting for Robot Localization
Mapping and localization are crucial problems in robotics and autonomous driving. Recent advances in 3D Gaussian Splatting (3DGS) have enabled precise 3D mapping and scene understanding by rendering photo-realistic image…
3DGSAutonomous DrivingScene UnderstandingWearable camera-based human absolute localization in large warehouses
In a robotised warehouse, as in any place where robots move autonomously, a major issue is the localization or detection of human operators during their intervention in the work area of the robots. This paper introduces …
Visual LocalizationDiaLoc: An Iterative Approach to Embodied Dialog Localization
Multimodal learning has advanced the performance for many vision-language tasks. However, most existing works in embodied dialog research focus on navigation and leave the localization task understudied. The few existing…
Unlocking the Capabilities of Vision-Language Models for Generalizable and Explainable Deepfake Detection
Current vision-language models (VLMs) have demonstrated remarkable capabilities in understanding multimodal data, but their potential remains underexplored for deepfake detection due to the misaligned of their knowledge …
Contrastive LearningDeepFake DetectionFace SwappingLarge Language ModelObject-Guided Day-Night Visual Localization in Urban Scenes
We introduce Object-Guided Localization (OGuL) based on a novel method of local-feature matching. Direct matching of local features is sensitive to significant changes in illumination. In contrast, object detection often…
Objectobject-detectionObject DetectionVisual Localization