paper-with-me

Papers

Adaptive Speech Quality Aware Complex Neural Network for Acoustic Echo Cancellation with Supervised Contrastive Learning

2022-10-30 · Bozhong Liu, Xiaoxi Yu, Hantao Huang

Acoustic echo cancellation (AEC) is designed to remove echoes, reverberation, and unwanted added sounds from the microphone signal while maintaining the quality of the near-end speaker's speech. This paper proposes adaptive speech quality complex neural networks to focus on specific tasks for real-time acoustic echo cancellation. In specific, we propose a complex modularize neural network with different stages to focus on feature extraction, acoustic separation, and mask optimization receptively. Furthermore, we adopt the contrastive learning framework and novel speech quality aware loss functions to further improve the performance. The model is trained with 72 hours for pre-training and then 72 hours for fine-tuning. The proposed model outperforms the state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2210.16791

Code (0)

등록된 구현이 없습니다.

Tasks

Acoustic echo cancellationContrastive Learning

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

A$^3$T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing

2022-03-18 · He Bai, Renjie Zheng, Junkun Chen, Xintong Li 외

Recently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation. However, all the above tasks are in the direction of spee…

Representation LearningSpeaker Verificationspeech-recognitionSpeech Recognition+5

AdaSpeech: Adaptive Text to Speech for Custom Voice

2021-03-01 · ICLR 2021 1 · Mingjian Chen, Xu Tan, Bohan Li, Yanqing Liu 외

Custom voice, a specific text to speech (TTS) service in commercial speech platforms, aims to adapt a source TTS model to synthesize personal voice for a target speaker using few speech data. Custom voice presents two un…

text-to-speechText to Speech

Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction

2024-01-03 · Minchan Kim, Myeonghun Jeong, Byoung Jin Choi, Semin Kim 외

We propose a novel text-to-speech (TTS) framework centered around a neural transducer. Our approach divides the whole TTS pipeline into semantic-level sequence-to-sequence (seq2seq) modeling and fine-grained acoustic mod…

text-to-speechText to Speech

StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization

2020-11-03 · Ahmed Mustafa, Nicola Pia, Guillaume Fuchs

In recent years, neural vocoders have surpassed classical speech generation approaches in naturalness and perceptual quality of the synthesized speech. Computationally heavy models like WaveNet and WaveGlow achieve best …

Spectral Reconstructiontext-to-speechText to SpeechVocal Bursts Intensity Prediction

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

2025-06-11 · Neta Glazer, Aviv Navon, Yael Segal, Aviv Shamsian 외

Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce UmbraTTS, a flow-matching based TTS mode…

Speech Synthesistext-to-speechText to Speech