paper-with-me

홈 › Papers

Play it by Ear: Learning Skills amidst Occlusion through Audio-Visual Imitation Learning

2022-05-30 · Maximilian Du, Olivia Y. Lee, Suraj Nair, Chelsea Finn

Humans are capable of completing a range of challenging manipulation tasks that require reasoning jointly over modalities such as vision, touch, and sound. Moreover, many such tasks are partially-observed; for example, taking a notebook out of a backpack will lead to visual occlusion and require reasoning over the history of audio or tactile information. While robust tactile sensing can be costly to capture on robots, microphones near or on a robot's gripper are a cheap and easy way to acquire audio feedback of contact events, which can be a surprisingly valuable data source for perception in the absence of vision. Motivated by the potential for sound to mitigate visual occlusion, we aim to learn a set of challenging partially-observed manipulation tasks from visual and audio inputs. Our proposed system learns these tasks by combining offline imitation learning from a modest number of tele-operated demonstrations and online finetuning using human provided interventions. In a set of simulated tasks, we find that our system benefits from using audio, and that by using online interventions we are able to improve the success rate of offline imitation learning by ~20%. Finally, we find that our system can complete a set of challenging, partially-observed tasks on a Franka Emika Panda robot, like extracting keys from a bag, with a 70% success rate, 50% higher than a policy that does not use audio.

📄 PDF Abstract BibTeX arXiv:2205.14850

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation Learning

Similar Papers 제목 키워드 기반

Mxplainer: Explain and Learn Insights by Imitating Mahjong Agents

2025-06-17 · Lingfeng li, Yunlong Lu, Yongyi Wang, Qifan Zheng 외

People need to internalize the skills of AI agents to improve their own capabilities. Our paper focuses on Mahjong, a multiplayer game involving imperfect information and requiring effective long-term decision-making ami…

Decision Making

Enhancing 3D Human Pose Estimation Amidst Severe Occlusion with Dual Transformer Fusion

2024-10-06 · Mehwish Ghafoor, Arif Mahmood, Muhammad Bilal

In the field of 3D Human Pose Estimation from monocular videos, the presence of diverse occlusion types presents a formidable challenge. Prior research has made progress by harnessing spatial and temporal cues to infer 3…

3D Human Pose Estimation3D Pose EstimationPose Estimation

My lips are concealed: Audio-visual speech enhancement through obstructions

2019-07-11 · Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman

Our objective is an audio-visual model for separating a single speaker from a mixture of sounds such as other speakers and background noise. Moreover, we wish to hear the speaker even when the visual cues are temporarily…

Speech Enhancement

TrueSkill Through Time: Revisiting the History of Chess

2017-12-03 · NIPS 2007 12 · Pierre Dangauthier, Ralf Herbrich, Tom Minka, Thore Graepel

We extend the Bayesian skill rating system TrueSkill to infer entire time series of skills of players by smoothing through time instead of filtering. The skill of each participating player, say, every year is represent…

Time SeriesTime Series Analysis

AIMusicGuru: Music Assisted Human Pose Correction

2022-03-24 · Snehesh Shrestha, Cornelia Fermüller, Tianyu Huang, Pyone Thant Win 외

Pose Estimation techniques rely on visual cues available through observations represented in the form of pixels. But the performance is bounded by the frame rate of the video and struggles from motion blur, occlusions, a…

Pose Estimation