paper-with-me

Papers

Speech Separation with Pretrained Frontend to Minimize Domain Mismatch

2024-11-05 · Wupeng Wang, Zexu Pan, Xinke Li, Shuai Wang, Haizhou Li

Speech separation seeks to separate individual speech signals from a speech mixture. Typically, most separation models are trained on synthetic data due to the unavailability of target reference in real-world cocktail party scenarios. As a result, there exists a domain gap between real and synthetic data when deploying speech separation models in real-world applications. In this paper, we propose a self-supervised domain-invariant pretrained (DIP) frontend that is exposed to mixture data without the need for target reference speech. The DIP frontend utilizes a Siamese network with two innovative pretext tasks, mixture predictive coding (MPC) and mixture invariant coding (MIC), to capture shared contextual cues between real and synthetic unlabeled mixtures. Subsequently, we freeze the DIP frontend as a feature extractor when training the downstream speech separation models on synthetic data. By pretraining the DIP frontend with the contextual cues, we expect that the speech separation skills learned from synthetic data can be effectively transferred to real data. To benefit from the DIP frontend, we introduce a novel separation pipeline to align the feature resolution of the separation models. We evaluate the speech separation quality on standard benchmarks and real-world datasets. The results confirm the superiority of our DIP frontend over existing speech separation models. This study underscores the potential of large-scale pretraining to enhance the quality and intelligibility of speech separation in real-world applications.

📄 PDF Abstract BibTeX arXiv:2411.03085

Code (1)

Wufan0Willan/DIP 공식 구현 pytorch

Tasks

Speech Separation

Methods 이 논문이 사용한 방법론

Siamese Network 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation

2025-04-03 · Wupeng Wang, Zexu Pan, Xinke Li, Shuai Wang 외

Speech separation (SS) seeks to disentangle a multi-talker speech mixture into single-talker speech streams. Although SS can be generally achieved using offline methods, such a processing paradigm is not suitable for rea…

DecoderKnowledge DistillationSpeech Separation

Bring the Noise: Introducing Noise Robustness to Pretrained Automatic Speech Recognition

2023-09-05 · Patrick Eickhoff, Matthias Möller, Theresa Pekarek Rosin, Johannes Twiefel 외

In recent research, in the domain of speech processing, large End-to-End (E2E) systems for Automatic Speech Recognition (ASR) have reported state-of-the-art performance on various benchmarks. These systems intrinsically …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDenoising+2

A Conformer-based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech Separation

2021-11-18 · Tom O'Malley, Arun Narayanan, Quan Wang, Alex Park 외

We present a frontend for improving robustness of automatic speech recognition (ASR), that jointly implements three modules within a single model: acoustic echo cancellation, speech enhancement, and speech separation. Th…

Acoustic echo cancellationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancement+3

End-to-End Multi-speaker ASR with Independent Vector Analysis

2022-04-01 · Robin Scheibler, Wangyou Zhang, Xuankai Chang, Shinji Watanabe 외

We develop an end-to-end system for multi-channel, multi-speaker automatic speech recognition. We propose a frontend for joint source separation and dereverberation based on the independent vector analysis (IVA) paradigm…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend

2021-02-23 · Wangyou Zhang, Christoph Boeddeker, Shinji Watanabe, Tomohiro Nakatani 외

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both single-channel and multichannel conditions. However, severe performance degradation is still obse…

Action DetectionActivity DetectionSpeech Dereverberationspeech-recognition+2