paper-with-me

Papers

Learning to Act with Affordance-Aware Multimodal Neural SLAM

2022-01-24 · Zhiwei Jia, Kaixiang Lin, Yizhou Zhao, Qiaozi Gao, Govind Thattai, Gaurav Sukhatme

Recent years have witnessed an emerging paradigm shift toward embodied artificial intelligence, in which an agent must learn to solve challenging tasks by interacting with its environment. There are several challenges in solving embodied multimodal tasks, including long-horizon planning, vision-and-language grounding, and efficient exploration. We focus on a critical bottleneck, namely the performance of planning and navigation. To tackle this challenge, we propose a Neural SLAM approach that, for the first time, utilizes several modalities for exploration, predicts an affordance-aware semantic map, and plans over it at the same time. This significantly improves exploration efficiency, leads to robust long-horizon planning, and enables effective vision-and-language grounding. With the proposed Affordance-aware Multimodal Neural SLAM (AMSLAM) approach, we obtain more than 40% improvement over prior published work on the ALFRED benchmark and set a new state-of-the-art generalization performance at a success rate of 23.48% on the test unseen scenes.

📄 PDF Abstract BibTeX arXiv:2201.09862

Code (1)

amazon-research/multimodal-neuralslam 공식 구현 pytorch

Tasks

Efficient ExplorationTest unseen

Similar Papers 제목 키워드 기반

Structure Aware SLAM using Quadrics and Planes

2018-04-24 · Mehdi Hosseinzadeh, Yasir Latif, Trung Pham, Niko Suenderhauf 외

Simultaneous Localization And Mapping (SLAM) is a fundamental problem in mobile robotics. While point-based SLAM methods provide accurate camera localization, the generated maps lack semantic information. On the other ha…

Camera LocalizationObjectobject-detectionObject Detection+1

Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation

2025-12-22 · Ryosuke Korekata, Quanting Xie, Yonatan Bisk, Komei Sugiura arxiv

In this study, we address the problem of open-vocabulary mobile manipulation, where a robot is required to carry a wide range of objects to receptacles based on free-form natural language instructions. This task is chall…

MAAL: Multimodality-Aware Autoencoder-Based Affordance Learning for 3D Articulated Objects

2023-01-01 · ICCV 2023 1 · Yuanzhi Liang, Xiaohan Wang, Linchao Zhu, Yi Yang

Inferring affordance for 3D articulated objects is a challenging and practical problem. It is a primary problem for applying robots to real-world scenarios. The exploration can be summarized as figuring out where to …

MMEObject

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

2026-02-10 · Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA 외 arxiv

3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for robotic manipulation. Previous research primarily focused on learning affordance knowledge from static cues such as lan…

Action UnderstandingPoint Clouds

RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation

2025-11-16 · Xiaoshuai Hao, Yingbo Tang, Lingfeng Zhang, Yanbiao Ma 외 arxiv

Robotic manipulation and navigation are fundamental capabilities of embodied intelligence, enabling effective robot interactions with the physical world. Achieving these capabilities requires a cohesive understanding of …

Affordance RecognitionScene UnderstandingObject RecognitionQuestion Answering