Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant
Imagine a future when we can Zoom-call a robot to manage household chores remotely. This work takes one step in this direction. Robi Butler is a new household robot assistant that enables seamless multimodal remote interaction. It allows the human user to monitor its environment from a first-person view, issue voice or text commands, and specify target objects through hand-pointing gestures. At its core, a high-level behavior module, powered by Large Language Models (LLMs), interprets multimodal instructions to generate multistep action plans. Each plan consists of open-vocabulary primitives supported by vision-language models, enabling the robot to process both textual and gestural inputs. Zoom provides a convenient interface to implement remote interactions between the human and the robot. The integration of these components allows Robi Butler to ground remote multimodal instructions in real-world home environments in a zero-shot manner. We evaluated the system on various household tasks, demonstrating its ability to execute complex user commands with multimodal inputs. We also conducted a user study to examine how multimodal interaction influences user experiences in remote human-robot interaction. These results suggest that with the advances in robot foundation models, we are moving closer to the reality of remote household robot assistants.
Code (0)
등록된 구현이 없습니다.
Tasks
multimodal interactionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
HoME: a Household Multimodal Environment
We introduce HoME: a Household Multimodal Environment for artificial agents to learn from vision, audio, semantics, physics, and interaction with objects and other agents, all within a realistic context. HoME integrates …
OpenAI Gymreinforcement-learningReinforcement LearningReinforcement Learning (RL)TokenButler: Token Importance is Predictable
Large Language Models (LLMs) rely on the Key-Value (KV) Cache to store token history, enabling efficient decoding of tokens. As the KV-Cache grows, it becomes a major memory and computation bottleneck, however, there is …
Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation
Given a natural language expression and a remote sensing image, the goal of referring remote sensing image segmentation (RRSIS) is to generate a pixel-level mask of the target object identified by the referring expressio…
BenchmarkingImage SegmentationReferring ExpressionSemantic SegmentationTracking Human Behavioural Consistency by Analysing Periodicity of Household Water Consumption
People are living longer than ever due to advances in healthcare, and this has prompted many healthcare providers to look towards remote patient care as a means to meet the needs of the future. It is now a priority to en…
Peoples Water Data: Enabling Reliable Field Data Generation and Microbial Contamination Screening in Household Drinking Water
Unsafe drinking water remains a major public health concern globally, particularly in low-resource regions where routine microbiological surveillance is limited. Although Escherichia coli is the internationally recognize…