paper-with-me

Papers

Egocentric Deep Multi-Channel Audio-Visual Active Speaker Localization

2022-01-06 · CVPR 2022 1 · Hao Jiang, Calvin Murdock, Vamsi Krishna Ithapu

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these social interactions first requires detecting and localizing the voice activities of the device wearer and the surrounding people. These tasks are challenging due to their egocentric nature: the wearer's head motion may cause motion blur, surrounding people may appear in difficult viewing angles, and there may be occlusions, visual clutter, audio noise, and bad lighting. Under these conditions, previous state-of-the-art active speaker detection methods do not give satisfactory results. Instead, we tackle the problem from a new setting using both video and multi-channel microphone array audio. We propose a novel end-to-end deep learning approach that is able to give robust voice activity detection and localization results. In contrast to previous methods, our method localizes active speakers from all possible directions on the sphere, even outside the camera's field of view, while simultaneously detecting the device wearer's own voice activity. Our experiments show that the proposed method gives superior results, can run in real time, and is robust against noise and clutter.

📄 PDF Abstract BibTeX arXiv:2201.01928

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActive Speaker DetectionActive Speaker LocalizationActivity DetectionAudio-Visual Active Speaker Detection

Similar Papers 제목 키워드 기반

Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos

2023-07-10 · CVPR 2024 1 · Sagnik Majumder, Ziad Al-Halah, Kristen Grauman

We propose a self-supervised method for learning representations based on spatial audio-visual correspondences in egocentric videos. Our method uses a masked auto-encoding framework to synthesize masked binaural (multi-c…

Active Speaker DetectionAudio DenoisingDenoising

Egocentric Audio-Visual Noise Suppression

2022-11-07 · Roshan Sharma, Weipeng He, Ju Lin, Egor Lakomkin 외

This paper studies audio-visual noise suppression for egocentric videos -- where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen s…

Action ClassificationEvent DetectionMulti-Task LearningObject Recognition+1

Chat2Map: Efficient Scene Mapping from Multi-Ego Conversations

2023-01-04 · CVPR 2023 1 · Sagnik Majumder, Hao Jiang, Pierre Moulon, Ethan Henderson 외

Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a prev…

Deep Reinforcement Learning

Spherical World-Locking for Audio-Visual Localization in Egocentric Videos

2024-08-09 · Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla, Anurag Kumar 외

Egocentric videos provide comprehensive contexts for user and scene understanding, spanning multisensory perception to behavioral interaction. We propose Spherical World-Locking (SWL) as a general framework for egocentri…

Active Speaker LocalizationDecoderScene UnderstandingVideo Understanding+1

Audio-Visual Talker Localization in Video for Spatial Sound Reproduction

2024-06-01 · Davide Berghi, Philip J. B. Jackson

Object-based audio production requires the positional metadata to be defined for each point-source object, including the key elements in the foreground of the sound scene. In many media production use cases, both cameras…

Active Speaker Detection