paper-with-me

Papers

LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music

2026-06-30 · Snehasis Banerjee, Ranjan Dasgupta arxiv

The quest for intuitive and natural human-robot interaction (HRI) remains a significant challenge in robotics. Traditional methods often rely on rigid, pre-programmed commands that limit the robot's expressiveness and adaptability. This paper introduces a novel framework that leverages the reasoning capabilities of Large Language Models (LLMs) to synthesize complex robotic actions from a rich tapestry of multimodal human inputs: natural speech, hand gestures, and music/sound beats. Our system architecture integrates a speech transcription model, a gesture recognition module, and a signal processing pipeline for beat detection. These processed inputs are contextualized using prompt templates and fed into a LLM. The LLM, informed by a predefined robot action space, reasons over the combined inputs to generate a coherent sequence of actions. This sequence is dispatched to an action queue for execution on a quadruped robot over ROS. The framework has ability to interpret and fuse semantic commands from speech, deictic information from gestures, and rhythmic cues from music. This work represents a step towards creating robots that can interact with humans in a more fluid, creative, and context-aware manner.

📄 PDF Abstract BibTeX arXiv:2606.31158

Code (0)

등록된 구현이 없습니다.

Tasks

Gesture Recognition

Similar Papers 제목 키워드 기반

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation

2025-12-01 · Ziqian Wang, Yonghao He, Licheng Yang, Wei Zou 외 arxiv

Simulation provides a low-cost, scalable pathway to large-scale robotic manipulation data collection. However, existing 3D scene generation methods can rarely be applied directly to manipulation data synthesis, as their …

Scene Generation

Chat-to-Design: AI Assisted Personalized Fashion Design

2022-07-03 · Weiming Zhuang, Chongjie Ye, Ying Xu, Pengzhi Mao 외

In this demo, we present Chat-to-Design, a new multimodal interaction system for personalized fashion design. Compared to classic systems that recommend apparel based on keywords, Chat-to-Design enables users to design c…

multimodal interactionNatural Language UnderstandingRetrieval

RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis

2024-02-25 · Yao Mu, Junting Chen, Qinglong Zhang, Shoufa Chen 외

Robotic behavior synthesis, the problem of understanding multimodal inputs and generating precise physical control for robots, is an important part of Embodied AI. Despite successes in applying multimodal large language …

Code GenerationMultimodal ReasoningVisual Question Answering

An interdisciplinary approach to high school curriculum development: Swarming Powered by Neuroscience

2021-09-12 · Elise Buckley, Joseph D. Monaco, Kevin M. Schultz, Robert Chalmers 외

This article discusses how to create an interactive virtual training program at the intersection of neuroscience, robotics, and computer science for high school students. A four-day microseminar, titled Swarming Powered …

AromaGen: Interactive Generation of Rich Olfactory Experiences with Multimodal Language Models

2026-04-02 · Yunge Wen, Awu Chen, Jianing Yu, Jas Brooks 외 arxiv

Smell's deep connection with food, memory, and social experience has long motivated researchers to bring olfaction into interactive systems. Yet most olfactory interfaces remain limited to fixed scent cartridges and pre-…