paper-with-me

홈 › Papers

Improving Perceptual Quality, Intelligibility, and Acoustics on VoIP Platforms

2023-03-16 · Joseph Konan, Ojas Bhargave, Shikhar Agnihotri, Hojeong Lee, Ankit Shah, Shuo Han, Yunyang Zeng, Amanda Shu, Haohui Liu, Xuankai Chang, Hamza Khalid, Minseon Gwak, Kawon Lee, Minjeong Kim, Bhiksha Raj

In this paper, we present a method for fine-tuning models trained on the Deep Noise Suppression (DNS) 2020 Challenge to improve their performance on Voice over Internet Protocol (VoIP) applications. Our approach involves adapting the DNS 2020 models to the specific acoustic characteristics of VoIP communications, which includes distortion and artifacts caused by compression, transmission, and platform-specific processing. To this end, we propose a multi-task learning framework for VoIP-DNS that jointly optimizes noise suppression and VoIP-specific acoustics for speech enhancement. We evaluate our approach on a diverse VoIP scenarios and show that it outperforms both industry performance and state-of-the-art methods for speech enhancement on VoIP applications. Our results demonstrate the potential of models trained on DNS-2020 to be improved and tailored to different VoIP platforms using VoIP-DNS, whose findings have important applications in areas such as speech recognition, voice assistants, and telecommunication.

📄 PDF Abstract BibTeX arXiv:2303.09048

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task LearningSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Psychoacoustic Challenges Of Speech Enhancement On VoIP Platforms

2023-10-11 · Joseph Konan, Shikhar Agnihotri, Ojas Bhargave, Shuo Han 외

Within the ambit of VoIP (Voice over Internet Protocol) telecommunications, the complexities introduced by acoustic transformations merit rigorous analysis. This research, rooted in the exploration of proprietary sender-…

BenchmarkingDenoisingSpeech Enhancement

Cellular Network Speech Enhancement: Removing Background and Transmission Noise

2023-01-22 · Amanda Shu, Hamza Khalid, Haohui Liu, Shikhar Agnihotri 외

The primary objective of speech enhancement is to reduce background noise while preserving the target's speech. A common dilemma occurs when a speaker is confined to a noisy environment and receives a call with high back…

Speech Enhancement

Dataset of British English speech recordings for psychoacoustics and speech processing research

2022-02-15 · Data in Brief 2022 2 · Trevor John Cox, Simone Graetzer, Michael A Akeroyd, Jonathan Barker 외

The Clarity Speech Corpus is a forty speaker British English speech dataset. The corpus was created for the purpose of running listening tests to gauge speech intelligibility and quality in the Clarity Project, which has…

Sentence

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

2025-12-24 · Haoyang Li, Xuyi Zhuang, Azmat Adnan, Ye Ni 외 arxiv

Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We propose GenTSE, a two-stage decoder-only generative …

V-Cloak: Intelligibility-, Naturalness- & Timbre-Preserving Real-Time Voice Anonymization

2022-10-27 · Jiangyi Deng, Fei Teng, Yanjiao Chen, Xiaofu Chen 외

Voice data generated on instant messaging or social media applications contains unique user voiceprints that may be abused by malicious adversaries for identity inference or identity theft. Existing voice anonymization t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Verificationspeech-recognition+2