paper-with-me

홈 › Papers

Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot Interaction

2026-06-09 · Dongjun Lee, Juheon Choi, Dong Kyu Shin, Sinjae Kang, Kimin Lee arxiv

For natural human-robot interaction, a robot must understand human intent expressed not only through language but also through nonverbal signals such as gestures and gaze. However, current robot policies rely on language instructions as the sole interface for conveying intent, leaving nonverbal signals unused and placing the full burden of communication. In this work, we present EDITH, a robot framework that captures the human's nonverbal signals through continuous streams of first-person view and gaze from smart glasses, and uses them alongside language instructions as inputs to the robot policy. Our hardware system streams the human's first-person view, gaze, and speech to the robot in real time, transcribing the speech into language instructions. To handle these rich but noisy signals, we design a hierarchical policy in which a high-level policy infers the human's intent and produces a sequence of subtasks, where each subtask is represented as a fine-grained instruction paired with a keyframe that grounds the intent in the scene (e.g., the frame where the human points at the target object). A low-level policy then executes these subtasks. In our experiments on human-robot interactive tasks, EDITH enables the robot to act on the human's nonverbal signals even when intent is expressed only briefly, and significantly reduces user effort to convey intent compared to using language instructions alone. Visit our project page for source code and real-robot demo videos.

📄 PDF Abstract BibTeX arXiv:2606.10276

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning

2026-06-16 · Taowen Wang, Zikang Xie, Bin Yang, Yunheng Wang 외 arxiv

Humanoid robots promise whole-body interaction in human-centered environments, but scalable policy learning remains difficult because task-level decision-making and whole-body dynamic execution are tightly coupled. A pra…

Decision Making

Episodic Memory Verbalization using Hierarchical Representations of Life-Long Robot Experience

2024-09-26 · Leonard Bärmann, Chad DeChant, Joana Plewnia, Fabian Peller-Konrad 외

Verbalization of robot experience, i.e., summarization of and question answering about a robot's past, is a crucial ability for improving human-robot interaction. Previous works applied rule-based systems or fine-tuned d…

Language ModelingLanguage ModellingLarge Language ModelQuestion Answering

EgoAdapt: Enhancing Robustness in Egocentric Interactive Speaker Detection Under Missing Modalities

2026-03-18 · Xinyuan Qian, Xinjia Zhu, Alessio Brutti, Dong Liang arxiv

TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-worl…

VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation

2025-09-24 · Shaofeng Yin, Yanjie Ze, Hong-Xing Yu, C. Karen Liu 외 arxiv

Humanoid loco-manipulation in unstructured environments demands tight integration of egocentric perception and whole-body control. However, existing approaches either depend on external motion capture systems or fail to …

Egocentric Action-aware Inertial Localization in Point Clouds

2025-05-20 · Mingfang Zhang, Ryo Yonetani, Yifei HUANG, Liangyang Ouyang 외

This paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages egocentric action cues from head-mounted IMU signals to localize the target individu…

Action Recognition