paper-with-me

홈 › Papers

Multimodal Human-Intent Modeling for Contextual Robot-to-Human Handovers of Arbitrary Objects

2025-08-05 · Lucas Chen, Guna Avula, Hanwen Ren, Zixing Wang, Ahmed H. Qureshi arxiv

Human-robot object handover is a crucial element for assistive robots that aim to help people in their daily lives, including elderly care, hospitals, and factory floors. The existing approaches to solving these tasks rely on pre-selected target objects and do not contextualize human implicit and explicit preferences for handover, limiting natural and smooth interaction between humans and robots. These preferences can be related to the target object selection from the cluttered environment and to the way the robot should grasp the selected object to facilitate desirable human grasping during handovers. Therefore, this paper presents a unified approach that selects target distant objects using human verbal and non-verbal commands and performs the handover operation by contextualizing human implicit and explicit preferences to generate robot grasps and compliant handover motion sequences. We evaluate our integrated framework and its components through real-world experiments and user studies with arbitrary daily-life objects. The results of these evaluations demonstrate the effectiveness of our proposed pipeline in handling object handover tasks by understanding human preferences. Our demonstration videos can be found at https://youtu.be/6z27B2INl-s.

📄 PDF Abstract BibTeX arXiv:2508.02982

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models

2026-04-27 · Hamed Rahimi, Clemence Grislain, Adrien Jacquet Cretides, Olivier Sigaud 외 arxiv

Improving the effectiveness of human-robot interaction requires social robots to accurately infer human goals through robust intention understanding. This challenge is particularly critical in multimodal settings, where …

RoboOmni: Proactive Robot Manipulation in Omni-modal Context

2025-10-27 · Siyin Wang, Jinlan Fu, Feihong Liu, Xinzhe He 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision-Language-Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rel…

Robot Manipulation

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

2026-05-08 · Jaeyoung Choi, Hyeondong Kim, Yujin Kim, Daehee Park arxiv

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task r…

IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction

2025-10-09 · Yandu Chen, Kefan Gu, Yuqing Wen, Yucheng Zhao 외 arxiv

Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general-purpose embodied intelligence. However, current SO…

Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning

2025-04-14 · Azizul Zahid, Jie Fan, Farong Wang, Ashton Dy 외

Understanding action correspondence between humans and robots is essential for evaluating alignment in decision-making, particularly in human-robot collaboration and imitation learning within unstructured environments. W…

Decision MakingImitation Learning