Egocentric Deep Multi-Channel Audio-Visual Active Speaker Localization
Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these social interactions first requires detecting and localizing the voice activities of the device wearer and the surrounding people. These tasks are challenging due to their egocentric nature: the wearer's head motion may cause motion blur, surrounding people may appear in difficult viewing angles, and there may be occlusions, visual clutter, audio noise, and bad lighting. Under these conditions, previous state-of-the-art active speaker detection methods do not give satisfactory results. Instead, we tackle the problem from a new setting using both video and multi-channel microphone array audio. We propose a novel end-to-end deep learning approach that is able to give robust voice activity detection and localization results. In contrast to previous methods, our method localizes active speakers from all possible directions on the sphere, even outside the camera's field of view, while simultaneously detecting the device wearer's own voice activity. Our experiments show that the proposed method gives superior results, can run in real time, and is robust against noise and clutter.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActive Speaker DetectionActive Speaker LocalizationActivity DetectionAudio-Visual Active Speaker DetectionSimilar Papers 제목 키워드 기반
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
We propose a self-supervised method for learning representations based on spatial audio-visual correspondences in egocentric videos. Our method uses a masked auto-encoding framework to synthesize masked binaural (multi-c…
Active Speaker DetectionAudio DenoisingDenoisingEgocentric Audio-Visual Noise Suppression
This paper studies audio-visual noise suppression for egocentric videos -- where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen s…
Action ClassificationEvent DetectionMulti-Task LearningObject Recognition+1Chat2Map: Efficient Scene Mapping from Multi-Ego Conversations
Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a prev…
Deep Reinforcement LearningSpherical World-Locking for Audio-Visual Localization in Egocentric Videos
Egocentric videos provide comprehensive contexts for user and scene understanding, spanning multisensory perception to behavioral interaction. We propose Spherical World-Locking (SWL) as a general framework for egocentri…
Active Speaker LocalizationDecoderScene UnderstandingVideo Understanding+1Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
Object-based audio production requires the positional metadata to be defined for each point-source object, including the key elements in the foreground of the sound scene. In many media production use cases, both cameras…
Active Speaker Detection