Memory-based gaze prediction in deep imitation learning for robot manipulation
Deep imitation learning is a promising approach that does not require hard-coded control rules in autonomous robot manipulation. The current applications of deep imitation learning to robot manipulation have been limited to reactive control based on the states at the current time step. However, future robots will also be required to solve tasks utilizing their memory obtained by experience in complicated environments (e.g., when the robot is asked to find a previously used object on a shelf). In such a situation, simple deep imitation learning may fail because of distractions caused by complicated environments. We propose that gaze prediction from sequential visual input enables the robot to perform a manipulation task that requires memory. The proposed algorithm uses a Transformer-based self-attention architecture for the gaze estimation based on sequential data to implement memory. The proposed method was evaluated with a real robot multi-object manipulation task that requires memory of the previous states.
Code (0)
등록된 구현이 없습니다.
Tasks
Gaze EstimationGaze PredictionImitation LearningRobot ManipulationSimilar Papers 제목 키워드 기반
Enhancing Reusability of Learned Skills for Robot Manipulation via Gaze and Bottleneck
Autonomous agents capable of diverse object manipulations should be able to acquire a wide range of manipulation skills with high reusability. Although advances in deep learning have made it increasingly feasible to repl…
Imitation LearningObjectRobot ManipulationGaze-based dual resolution deep imitation learning for high-precision dexterous robot manipulation
A high-precision manipulation task, such as needle threading, is challenging. Physiological studies have proposed connecting low-resolution peripheral vision and fast movement to transport the hand into the vicinity of a…
Computational EfficiencyImitation LearningRobot ManipulationGaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent.…
Robot ManipulationIntent at a Glance: Gaze-Guided Robotic Manipulation via Foundation Models
Designing intuitive interfaces for robotic control remains a central challenge in enabling effective human-robot interaction, particularly in assistive care settings. Eye gaze offers a fast, non-intrusive, and intent-ric…
Intention estimation from gaze and motion features for human-robot shared-control object manipulation
Shared control can help in teleoperated object manipulation by assisting with the execution of the user's intention. To this end, robust and prompt intention estimation is needed, which relies on behavioral observations.…