paper-with-me

홈 › Papers

Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models

2025-01-07 · Kazuhiro Sasabuchi, Naoki Wake, Atsushi Kanehira, Jun Takamatsu, Katsushi Ikeuchi

In human-robot interaction (HRI), the beginning of an interaction is often complex. Whether the robot should communicate with the human is dependent on several situational factors (e.g., the current human's activity, urgency of the interaction, etc.). We test whether large language models (LLM) and vision language models (VLM) can provide solutions to this problem. We compare four different system-design patterns using LLMs and VLMs, and test on a test set containing 84 human-robot situations. The test set mixes several publicly available datasets and also includes situations where the appropriate action to take is open-ended. Our results using the GPT-4o and Phi-3 Vision model indicate that LLMs and VLMs are capable of handling interaction beginnings when the desired actions are clear, however, challenge remains in the open-ended situations where the model must balance between the human and robot situation.

📄 PDF Abstract BibTeX arXiv:2503.15491

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

When and How to Express Empathy in Human-Robot Interaction Scenarios

2025-09-11 · Christian Arzate Cruz, Edwin C. Montiel-Vazquez, Chikara Maeda, Randy Gomez arxiv

Incorporating empathetic behavior into robots can improve their social effectiveness and interaction quality. In this paper, we present whEE (when and how to express empathy), a framework that enables social robots to de…

Incremental Learning of Humanoid Robot Behavior from Natural Interaction and Large Language Models

2023-09-08 · Leonard Bärmann, Rainer Kartmann, Fabian Peller-Konrad, Jan Niehues 외

Natural-language dialog is key for intuitive human-robot interaction. It can be used not only to express humans' intents, but also to communicate instructions for improvement if a robot does not understand a command corr…

Incremental LearningPrompt Learning

A Multimodal Framework for Human-Multi-Agent Interaction

2026-03-24 · Shaid Hasan, Breenice Lee, Sujan Sarker, Tariq Iqbal arxiv

Human-robot interaction is increasingly moving toward multi-robot, socially grounded environments. Existing systems struggle to integrate multimodal perception, embodied expression, and coordinated decision-making in a u…

Multimodal Reasoning

A Sign Language Recognition System with Pepper, Lightweight-Transformer, and LLM

2023-09-28 · JongYoon Lim, Inkyu Sa, Bruce MacDonald, Ho Seok Ahn

This research explores using lightweight deep neural network architectures to enable the humanoid robot Pepper to understand American Sign Language (ASL) and facilitate non-verbal human-robot interaction. First, we intro…

Prompt EngineeringSign Language Recognition

A MultiModal Social Robot Toward Personalized Emotion Interaction

2021-10-08 · Baijun Xie, Chung Hyuk Park

Human emotions are expressed through multiple modalities, including verbal and non-verbal information. Moreover, the affective states of human users can be the indicator for the level of engagement and successful interac…

reinforcement-learningReinforcement Learning (RL)