paper-with-me

홈 › Papers

Robust Active Speaker Detection in Noisy Environments

2024-03-27 · Siva Sai Nagender Vasireddy, Chenxu Zhang, Xiaohu Guo, Yapeng Tian

This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, but non-speech sounds in the surrounding environment can negatively impact performance. To overcome this, we propose a novel framework that utilizes audio-visual speech separation as guidance to learn noise-free audio features. These features are then utilized in an ASD model, and both tasks are jointly optimized in an end-to-end framework. Our proposed framework mitigates residual noise and audio quality reduction issues that can occur in a naive cascaded two-stage framework that directly uses separated speech for ASD, and enables the two tasks to be optimized simultaneously. To further enhance the robustness of the audio features and handle inherent speech noises, we propose a dynamic weighted loss approach to train the speech separator. We also collected a real-world noise audio dataset to facilitate investigations. Experiments demonstrate that non-speech audio noises significantly impact ASD models, and our proposed approach improves ASD performance in noisy environments. The framework is general and can be applied to different ASD approaches to improve their robustness. Our code, models, and data will be released.

📄 PDF Abstract BibTeX arXiv:2403.19002

Code (0)

등록된 구현이 없습니다.

Tasks

Active Speaker DetectionSpeech Separation

Similar Papers 제목 키워드 기반

LSTM-CNN Network for Audio Signature Analysis in Noisy Environments

2023-12-12 · Praveen Damacharla, Hamid Rajabalipanah, Mohammad Hosein Fakheri

There are multiple applications to automatically count people and specify their gender at work, exhibitions, malls, sales, and industrial usage. Although current speech detection methods are supposed to operate well, in …

Best of Both Worlds: Multi-task Audio-Visual Automatic Speech Recognition and Active Speaker Detection

2022-05-10 · Otavio Braga, Olivier Siohan

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker's face. However, when multiple candidate speakers are visible this tra…

Active Speaker DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1

AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection

2019-01-05 · Joseph Roth, Sourish Chaudhuri, Ondrej Klejch, Radhika Marvin 외

Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence …

Active Speaker DetectionAudio-Visual Active Speaker DetectionDiversityspeaker-diarization+2

Audio-Visual Talker Localization in Video for Spatial Sound Reproduction

2024-06-01 · Davide Berghi, Philip J. B. Jackson

Object-based audio production requires the positional metadata to be defined for each point-source object, including the key elements in the foreground of the sound scene. In many media production use cases, both cameras…

Active Speaker Detection

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining

2025-01-06 · Holger Severin Bovbjerg, Jan Østergaard, Jesper Jensen, Zheng-Hua Tan

Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting the presence of speech from a known target-speaker in an audio frame. Recently, deep neural network-based models have shown good performance in th…

Action DetectionActivity DetectionDenoisingSelf-Supervised Learning