paper-with-me

홈 › Papers

CochleaNet: A Robust Language-independent Audio-Visual Model for Speech Enhancement

2019-09-23 · Mandar Gogate, Kia Dashtipour, Ahsan Adeel, Amir Hussain

Noisy situations cause huge problems for suffers of hearing loss as hearing aids often make the signal more audible but do not always restore the intelligibility. In noisy settings, humans routinely exploit the audio-visual (AV) nature of the speech to selectively suppress the background noise and to focus on the target speaker. In this paper, we present a causal, language, noise and speaker independent AV deep neural network (DNN) architecture for speech enhancement (SE). The model exploits the noisy acoustic cues and noise robust visual cues to focus on the desired speaker and improve the speech intelligibility. To evaluate the proposed SE framework a first of its kind AV binaural speech corpus, called ASPIRE, is recorded in real noisy environments including cafeteria and restaurant. We demonstrate superior performance of our approach in terms of objective measures and subjective listening tests over the state-of-the-art SE approaches as well as recent DNN based SE models. In addition, our work challenges a popular belief that a scarcity of multi-language large vocabulary AV corpus and wide variety of noises is a major bottleneck to build a robust language, speaker and noise independent SE systems. We show that a model trained on synthetic mixture of Grid corpus (with 33 speakers and a small English vocabulary) and ChiME 3 Noises (consisting of only bus, pedestrian, cafeteria, and street noises) generalise well not only on large vocabulary corpora but also on completely unrelated languages (such as Mandarin), wide variety of speakers and noises.

📄 PDF Abstract BibTeX arXiv:1909.10407

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

AV Speech Enhancement Challenge using a Real Noisy Corpus

2019-09-30 · Mandar Gogate, Ahsan Adeel, Kia Dashtipour, Peter Derleth 외

This paper presents, a first of its kind, audio-visual (AV) speech enhacement challenge in real-noisy settings. A detailed description of the AV challenge, a novel real noisy AV corpus (ASPIRE), benchmark speech enhancem…

Speech Enhancement

Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation

2018-04-10 · Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel 외

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and do…

Speech Separation

Audiovisual Speech Synthesis using Tacotron2

2020-08-03 · Ahmed Hussen Abdelaziz, Anushree Prasanna Kumar, Chloe Seivwright, Gabriele Fanelli 외

Audiovisual speech synthesis is the problem of synthesizing a talking face while maximizing the coherency of the acoustic and visual speech. In this paper, we propose and compare two audiovisual speech synthesis systems …

Face ModelSentenceSpeech Synthesis

VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning

2022-11-21 · Qiushi Zhu, Long Zhou, Ziqiang Zhang, Shujie Liu 외

Although speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework t…

Audio-Visual Speech RecognitionLanguage ModellingRepresentation Learningspeech-recognition+3

Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes

2025-03-24 · Hyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda Senocak

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches ar…

Cross-Modal RetrievalDisentanglementVisual Grounding