paper-with-me

홈 › Papers

An Approach to Combining Video and Speech with Large Language Models in Human-Robot Interaction

2026-02-23 · Guanting Shen, Zi Tian arxiv

Interpreting human intent accurately is a central challenge in human-robot interaction (HRI) and a key requirement for achieving more natural and intuitive collaboration between humans and machines. This work presents a novel multimodal HRI framework that combines advanced vision-language models, speech processing, and fuzzy logic to enable precise and adaptive control of a Dobot Magician robotic arm. The proposed system integrates Florence-2 for object detection, Llama 3.1 for natural language understanding, and Whisper for speech recognition, providing users with a seamless and intuitive interface for object manipulation through spoken commands. By jointly addressing scene perception and action planning, the approach enhances the reliability of command interpretation and execution. Experimental evaluations conducted on consumer-grade hardware demonstrate a command execution accuracy of 75\%, highlighting both the robustness and adaptability of the system. Beyond its current performance, the proposed architecture serves as a flexible and extensible foundation for future HRI research, offering a practical pathway toward more sophisticated and natural human-robot collaboration through tightly coupled speech and vision-language processing.

📄 PDF Abstract BibTeX arXiv:2602.20219

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language UnderstandingSpeech RecognitionObject Detection

Similar Papers 제목 키워드 기반

V-SAT: Video Subtitle Annotation Tool

2025-10-28 · Arpita Kundu, Joyita Chakraborty, Anindita Desarkar, Aritra Sen 외 arxiv

The surge of audiovisual content on streaming platforms and social media has heightened the demand for accurate and accessible subtitles. However, existing subtitle generation methods primarily speech-based transcription…

Speech Recognition

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

2024-06-22 · Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen 외

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end a…

DiversityLanguage ModelingLanguage ModellingLarge Language Model+1

Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge

2023-08-18 · ICCV 2023 1 · Minsu Kim, Jeong Hun Yeo, Jeongsoo Choi, Yong Man Ro

This paper proposes a novel lip reading framework, especially for low-resource languages, which has not been well addressed in the previous literature. Since low-resource languages do not have enough video-text paired da…

Lip Reading

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

2025-10-08 · Jiaben Chen, Zixin Wang, Ailing Zeng, Yang Fu 외 arxiv

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offer…

Video Generation

Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech

2024-09-23 · Hong Nguyen, Sean Foley, Kevin Huang, Xuan Shi 외

Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce…