paper-with-me

Papers

Deep Learning Based Audio-Visual Multi-Speaker DOA Estimation Using Permutation-Free Loss Function

2022-10-26 · Qing Wang, Hang Chen, Ya Jiang, Zhe Wang, Yuyang Wang, Jun Du, Chin-Hui Lee

In this paper, we propose a deep learning based multi-speaker direction of arrival (DOA) estimation with audio and visual signals by using permutation-free loss function. We first collect a data set for multi-modal sound source localization (SSL) where both audio and visual signals are recorded in real-life home TV scenarios. Then we propose a novel spatial annotation method to produce the ground truth of DOA for each speaker with the video data by transformation between camera coordinate and pixel coordinate according to the pin-hole camera model. With spatial location information served as another input along with acoustic feature, multi-speaker DOA estimation could be solved as a classification task of active speaker detection. Label permutation problem in multi-speaker related tasks will be addressed since the locations of each speaker are used as input. Experiments conducted on both simulated data and real data show that the proposed audio-visual DOA estimation model outperforms audio-only DOA estimation model by a large margin.

📄 PDF Abstract BibTeX arXiv:2210.14581

Code (0)

등록된 구현이 없습니다.

Tasks

Active Speaker DetectionSound Source Localization

Similar Papers 제목 키워드 기반

Audio-visual Speech Separation with Adversarially Disentangled Visual Representation

2020-11-29 · Peng Zhang, Jiaming Xu, Jing Shi, Yunzhe Hao 외

Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they build on a strategy to handle the predefin…

Speech Separation

FaceFilter: Audio-visual speech separation using still images

2020-05-14 · Soo-Whan Chung, Soyeon Choe, Joon Son Chung, Hong-Goo Kang

The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network. Unlike previous works that used lip movement on video clips or pre-…

Speech Separation

DNN driven Speaker Independent Audio-Visual Mask Estimation for Speech Separation

2018-07-31 · Mandar Gogate, Ahsan Adeel, Ricard Marxer, Jon Barker 외

Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues …

Speech Separation

Late Audio-Visual Fusion for In-The-Wild Speaker Diarization

2022-11-02 · Zexu Pan, Gordon Wichern, François G. Germain, Aswin Subramanian 외

Speaker diarization is well studied for constrained audios but little explored for challenging in-the-wild videos, which have more speakers, shorter utterances, and inconsistent on-screen speakers. We address this gap by…

speaker-diarizationSpeaker DiarizationSpeaker Recognition

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

2025-05-20 · Ming Gao, Shilong Wu, Hang Chen, Jun Du 외

Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on mul…

Audio-Visual Speech Recognitionspeaker-diarizationSpeaker Diarizationspeech-recognition+2