paper-with-me

홈 › Papers

WALL-E: Embodied Robotic WAiter Load Lifting with Large Language Model

2023-08-30 · Tianyu Wang, YiFan Li, Haitao Lin, xiangyang xue, Yanwei Fu

Enabling robots to understand language instructions and react accordingly to visual perception has been a long-standing goal in the robotics research community. Achieving this goal requires cutting-edge advances in natural language processing, computer vision, and robotics engineering. Thus, this paper mainly investigates the potential of integrating the most recent Large Language Models (LLMs) and existing visual grounding and robotic grasping system to enhance the effectiveness of the human-robot interaction. We introduce the WALL-E (Embodied Robotic WAiter load lifting with Large Language model) as an example of this integration. The system utilizes the LLM of ChatGPT to summarize the preference object of the users as a target instruction via the multi-round interactive dialogue. The target instruction is then forwarded to a visual grounding system for object pose and size estimation, following which the robot grasps the object accordingly. We deploy this LLM-empowered system on the physical robot to provide a more user-friendly interface for the instruction-guided grasping task. The further experimental results on various real-world scenarios demonstrated the feasibility and efficacy of our proposed framework. See the project website at: https://star-uu-wang.github.io/WALL-E/

📄 PDF Abstract BibTeX arXiv:2308.15962

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelObjectRobotic GraspingVisual Grounding

Similar Papers 제목 키워드 기반

Trajectory Planning of a Curtain Wall Installation Robot Based on Biomimetic Mechanisms

2025-07-22 · Xiao Liu, Weijun Wang, Tianlun Huang, Zhiyong Wang 외 arxiv

As the robotics market rapidly evolves, energy consumption has become a critical issue, particularly restricting the application of construction robots. To tackle this challenge, our study innovatively draws inspiration …

Trajectory Planning

Robotic Nonprehensile Object Transportation with a Hanging Tray

2026-06-08 · Adam Heins, Angela P. Schoellig arxiv

We consider the nonprehensile object transportation task known as the waiter's problem, in which a robot must move an object balanced on a tray from one location to another. In contrast to prior works on the robotic wait…

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics

2025-07-29 · Leonard Hinckeldey, Elliot Fosong, Rimvydas Rubavicius, Elle Miller 외 arxiv

The development of reinforcement learning (RL) algorithms has been largely driven by ambitious challenge tasks and benchmarks. Games have dominated RL benchmarks because they present relevant challenges, are inexpensive …

Reinforcement LearningContinuous Control

Vector Quantized Feature Fields for Fast 3D Semantic Lifting

2025-03-09 · George Tang, Aditya Agarwal, Weiqiao Han, Trevor Darrell 외

We generalize lifting to semantic lifting by incorporating per-view masks that indicate relevant pixels for lifting tasks. These masks are determined by querying corresponding multiscale pixel-aligned feature maps, which…

Embodied Question AnsweringQuestion Answering

Multimodal Spiking Neural Network for Space Robotic Manipulation

2025-08-10 · Liwen Zhang, Dong Zhou, Shibo Shao, Zihao Su 외 arxiv

This paper presents a multimodal control framework based on spiking neural networks (SNNs) for robotic arms aboard space stations. It is designed to cope with the constraints of limited onboard resources while enabling a…

Reinforcement Learning