paper-with-me

홈 › Papers

Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning

2025-04-14 · Azizul Zahid, Jie Fan, Farong Wang, Ashton Dy, Sai Swaminathan, Fei Liu

Understanding action correspondence between humans and robots is essential for evaluating alignment in decision-making, particularly in human-robot collaboration and imitation learning within unstructured environments. We propose a multimodal demonstration learning framework that explicitly models human demonstrations from RGB video with robot demonstrations in voxelized RGB-D space. Focusing on the "pick and place" task from the RH20T dataset, we utilize data from 5 users across 10 diverse scenes. Our approach combines ResNet-based visual encoding for human intention modeling and a Perceiver Transformer for voxel-based robot action prediction. After 2000 training epochs, the human model reaches 71.67% accuracy, and the robot model achieves 71.8% accuracy, demonstrating the framework's potential for aligning complex, multimodal human and robot behaviors in manipulation tasks.

📄 PDF Abstract BibTeX arXiv:2504.11493

Code (1)

utkauraslab/aligning_hr_actions 공식 구현 pytorch

Tasks

Decision MakingImitation Learning

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Residual Connection 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Bridging Speech, Emotion, and Motion: a VLM-based Multimodal Edge-deployable Framework for Humanoid Robots

2026-02-07 · Songhua Yang, Xuetao Li, Xuanye Fei, Mengde Li 외 arxiv

Effective human-robot interaction requires emotionally rich multimodal expressions, yet most humanoid robots lack coordinated speech, facial expressions, and gestures. Meanwhile, real-world deployment demands on-device s…

Expanding Frozen Vision-Language Models without Retraining: Towards Improved Robot Perception

2023-08-31 · Riley Tavassoli, Mani Amani, Reza Akhavian

Vision-language models (VLMs) have shown powerful capabilities in visual question answering and reasoning tasks by combining visual representations with the abstract skill set large language models (LLMs) learn during pr…

Activity RecognitionHuman Activity RecognitionQuestion AnsweringScene Understanding+1

Comparison of Lexical Alignment with a Teachable Robot in Human-Robot and Human-Human-Robot Interactions

2022-09-23 · SIGDIAL (ACL) 2022 9 · Yuya Asano, Diane Litman, Mingzhi Yu, Nikki Lobczowski 외

Speakers build rapport in the process of aligning conversational behaviors with each other. Rapport engendered with a teachable agent while instructing domain material has been shown to promote learning. Past work on lex…

The AICO Multimodal Corpus -- Data Collection and Preliminary Analyses

2020-05-01 · LREC 2020 5 · Kristiina Jokinen

This paper describes data collection and the first explorative research on the AICO Multimodal Corpus. The corpus contains eye-gaze, Kinect, and video recordings of human-robot and human-human interactions, and was colle…

iCub! Do you recognize what I am doing?: multimodal human action recognition on multisensory-enabled iCub robot

2022-12-17 · Kas Kniesmeijer, Murat Kirtay

This study uses multisensory data (i.e., color and depth) to recognize human actions in the context of multimodal human-robot interaction. Here we employed the iCub robot to observe the predefined actions of the human pa…

Action RecognitionEnsemble LearningTemporal Action Localization