paper-with-me

Papers

A Framework for Low-Latency, LLM-driven Multimodal Interaction on the Pepper Robot

2026-01-09 · Erich Studerus, Vivienne Jia Zhong, Stephan Vonschallen arxiv

Despite recent advances in integrating Large Language Models (LLMs) into social robotics, two weaknesses persist. First, existing implementations on platforms like Pepper often rely on cascaded Speech-to-Text (STT)->LLM->Text-to-Speech (TTS) pipelines, resulting in high latency and the loss of paralinguistic information. Second, most implementations fail to fully leverage the LLM's capabilities for multimodal perception and agentic control. We present an open-source Android framework for the Pepper robot that addresses these limitations through two key innovations. First, we integrate end-to-end Speech-to-Speech (S2S) models to achieve low-latency interaction while preserving paralinguistic cues and enabling adaptive intonation. Second, we implement extensive Function Calling capabilities that elevate the LLM to an agentic planner, orchestrating robot actions (navigation, gaze control, tablet interaction) and integrating diverse multimodal feedback (vision, touch, system state). The framework runs on the robot's tablet but can also be built to run on regular Android smartphones or tablets, decoupling development from robot hardware. This work provides the HRI community with a practical, extensible platform for exploring advanced LLM-driven embodied interaction.

📄 PDF Abstract BibTeX arXiv:2603.21013

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Database of Multimodal Data to Construct a Simulated Dialogue Partner with Varying Degrees of Cognitive Health

2022-06-01 · RaPID (LREC) 2022 6 · Ruihao Pan, Ziming Liu, Fengpei Yuan, Maryam Zare 외

An assistive robot that could communicate with dementia patients would have great social benefit. An assistive robot Pepper has been designed to administer Referential Communication Tasks (RCTs) to human subjects without…

Dialogue ManagementManagementReinforcement Learning (RL)

A Sign Language Recognition System with Pepper, Lightweight-Transformer, and LLM

2023-09-28 · JongYoon Lim, Inkyu Sa, Bruce MacDonald, Ho Seok Ahn

This research explores using lightweight deep neural network architectures to enable the humanoid robot Pepper to understand American Sign Language (ASL) and facilitate non-verbal human-robot interaction. First, we intro…

Prompt EngineeringSign Language Recognition

Forecasting the Number of Harvest-ready Fruits of Sweet Peppers Using Multimodal Time-Series Data

2026-07-22 · Enrico Pallotta, Mohamed Farag, Esra Guclu, Chris McCool 외 arxiv

Accurate yield forecasting at the individual-plant level is critical for precision agriculture and supply-chain planning, yet public datasets capturing both visual growth dynamics and per-plant measurement labels are sca…

Multimodal Deep Learning

Sequential annotations for naturally-occurring HRI: first insights

2023-08-29 · Lucien Tisserand, Frédéric Armetta, Heike Baldauf-Quilliatre, Antoine Bouquin 외

We explain the methodology we developed for improving the interactions accomplished by an embedded conversational agent, drawing from Conversation Analytic sequential and multimodal analysis. The use case is a Pepper rob…

SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters

2025-01-01 · CVPR 2025 1 · Jianping Jiang, Weiye Xiao, Zhengyu Lin, Huaizhong Zhang 외

Human beings are social animals. How to equip 3D autonomous characters with similar social intelligence that can perceive, understand and interact with humans remains an open yet foundamental problem. In this paper, …

Vision-Language-Action