paper-with-me

Papers

Self-Supervised Learning-Based Source Separation for Meeting Data

2023-04-03 · Yuang Li, Xianrui Zheng, Philip C. Woodland

Source separation can improve automatic speech recognition (ASR) under multi-party meeting scenarios by extracting single-speaker signals from overlapped speech. Despite the success of self-supervised learning models in single-channel source separation, most studies have focused on simulated setups. In this paper, seven SSL models were compared on both simulated and real-world corpora. Then, we propose to integrate the best-performing model WavLM into an automatic transcription system through a novel iterative source selection method. To improve real-world performance, time-domain unsupervised mixture invariant training was adapted to the time-frequency domain. Experiments showed that in the transcription system when source separation was inserted before an ASR model fine-tuned on separated speech, absolute reductions of 1.9% and 1.5% in concatenated minimum-permutation word error rate for an unknown number of speakers (cpWER-us) were observed on the AMI dev and test sets.

📄 PDF Abstract BibTeX arXiv:2304.00871

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models

2024-10-28 · Tobias Cord-Landwehr, Christoph Boeddeker, Reinhold Haeb-Umbach

We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separation, and a von-Mises-Fisher Mixture Mo…

Speech Enhancement

Speech separation with large-scale self-supervised learning

2022-11-09 · Zhuo Chen, Naoyuki Kanda, Jian Wu, Yu Wu 외

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the exploration of the SSL-based SS by massive…

Self-Supervised LearningSpeech Separation

PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings

2024-03-04 · Joonas Kalda, Clément Pagés, Ricard Marxer, Tanel Alumäe 외

A major drawback of supervised speech separation (SSep) systems is their reliance on synthetic data, leading to poor real-world generalization. Mixture invariant training (MixIT) was proposed as an unsupervised alternati…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+3

Pac-HuBERT: Self-Supervised Music Source Separation via Primitive Auditory Clustering and Hidden-Unit BERT

2023-04-04 · Ke Chen, Gordon Wichern, François G. Germain, Jonathan Le Roux

In spite of the progress in music source separation research, the small amount of publicly-available clean source data remains a constant limiting factor for performance. Thus, recent advances in self-supervised learning…

ClusteringDecoderMusic Source SeparationSelf-Supervised Learning

GPU-accelerated Guided Source Separation for Meeting Transcription

2022-12-10 · Desh Raj, Daniel Povey, Sanjeev Khudanpur

Guided source separation (GSS) is a type of target-speaker extraction method that relies on pre-computed speaker activities and blind source separation to perform front-end enhancement of overlapped speech signals. It wa…

blind source separationCPUGPUSpeech Recognition+1