paper-with-me

Papers

Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant

2024-09-30 · Anxing Xiao, Nuwan Janaka, Tianrun Hu, Anshul Gupta, Kaixin Li, Cunjun Yu, David Hsu

Imagine a future when we can Zoom-call a robot to manage household chores remotely. This work takes one step in this direction. Robi Butler is a new household robot assistant that enables seamless multimodal remote interaction. It allows the human user to monitor its environment from a first-person view, issue voice or text commands, and specify target objects through hand-pointing gestures. At its core, a high-level behavior module, powered by Large Language Models (LLMs), interprets multimodal instructions to generate multistep action plans. Each plan consists of open-vocabulary primitives supported by vision-language models, enabling the robot to process both textual and gestural inputs. Zoom provides a convenient interface to implement remote interactions between the human and the robot. The integration of these components allows Robi Butler to ground remote multimodal instructions in real-world home environments in a zero-shot manner. We evaluated the system on various household tasks, demonstrating its ability to execute complex user commands with multimodal inputs. We also conducted a user study to examine how multimodal interaction influences user experiences in remote human-robot interaction. These results suggest that with the advances in robot foundation models, we are moving closer to the reality of remote household robot assistants.

📄 PDF Abstract BibTeX arXiv:2409.20548

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal interaction

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

HoME: a Household Multimodal Environment

2017-11-29 · Simon Brodeur, Ethan Perez, Ankesh Anand, Florian Golemo 외

We introduce HoME: a Household Multimodal Environment for artificial agents to learn from vision, audio, semantics, physics, and interaction with objects and other agents, all within a realistic context. HoME integrates …

OpenAI Gymreinforcement-learningReinforcement LearningReinforcement Learning (RL)

TokenButler: Token Importance is Predictable

2025-03-10 · Yash Akhauri, Ahmed F AbouElhamayed, YiFei Gao, Chi-Chih Chang 외

Large Language Models (LLMs) rely on the Key-Value (KV) Cache to store token history, enabling efficient decoding of tokens. As the KV-Cache grows, it becomes a major memory and computation bottleneck, however, there is …

Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation

2024-10-11 · Zhe Dong, Yuzhe Sun, Yanfeng Gu, Tianzhu Liu

Given a natural language expression and a remote sensing image, the goal of referring remote sensing image segmentation (RRSIS) is to generate a pixel-level mask of the target object identified by the referring expressio…

BenchmarkingImage SegmentationReferring ExpressionSemantic Segmentation

Tracking Human Behavioural Consistency by Analysing Periodicity of Household Water Consumption

2019-10-16

People are living longer than ever due to advances in healthcare, and this has prompted many healthcare providers to look towards remote patient care as a means to meet the needs of the future. It is now a priority to en…

Peoples Water Data: Enabling Reliable Field Data Generation and Microbial Contamination Screening in Household Drinking Water

2026-04-05 · Suzan Kagan, Shira Spigelman, Sankar Sudhir, Thalappil Pradeep 외 arxiv

Unsafe drinking water remains a major public health concern globally, particularly in low-resource regions where routine microbiological surveillance is limited. Although Escherichia coli is the internationally recognize…