paper-with-me

홈 › Papers

SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR

2024-12-07 · Pengcheng Guo, Xuankai Chang, Hang Lv, Shinji Watanabe, Lei Xie

Benefiting from massive and diverse data sources, speech foundation models exhibit strong generalization and knowledge transfer capabilities to a wide range of downstream tasks. However, a limitation arises from their exclusive handling of single-speaker speech input, making them ineffective in recognizing multi-speaker overlapped speech, a common occurrence in real-world scenarios. In this study, we delve into the adaptation of speech foundation models to eliminate interfering speakers from overlapping speech and perform target-speaker automatic speech recognition (TS-ASR). Initially, we utilize the Whisper model as the foundation for adaptation and conduct a thorough comparison of its integration with existing target-speaker adaptation techniques. We then propose an innovative model termed Speaker-Querying Whisper (SQ-Whisper), which employs a set number of trainable queries to capture speaker prompts from overlapping speech based on target-speaker enrollment. These prompts serve to steer the model in extracting speaker-specific features and accurately recognizing target-speaker transcriptions. Experimental results demonstrate that our approach effectively adapts the pre-trained speech foundation model to TS-ASR. Compared with the robust TS-HuBERT model, the proposed SQ-Whisper significantly improves performance, yielding up to 15% and 10% relative reductions in word error rates (WERs) on the Libri2Mix and WSJ0-2Mix datasets, respectively. With data augmentation, we establish new state-of-the-art WERs of 14.6% on the Libri2Mix Test set and 4.4% on the WSJ0-2Mix Test set. Furthermore, we evaluate our model on the real-world AMI meeting dataset, which shows consistent improvement over other adaptation methods.

📄 PDF Abstract BibTeX arXiv:2412.05589

Code (1)

pengchengguo/espnet 공식 구현 pytorch

Tasks

Automatic Speech RecognitionData Augmentationspeech-recognitionSpeech RecognitionTransfer Learning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition

2024-12-30 · Alexander Polok, Dominik Klement, Martin Kocour, Jiangyu Han 외

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a significant challenge, particularly when systems conditioned on speaker embeddings fail to generalize to unseen speakers. In t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2

Voice Conversion for Whispered Speech Synthesis

2019-12-11 · Marius Cotescu, Thomas Drugman, Goeric Huybrechts, Jaime Lorenzo-Trueba 외

We present an approach to synthesize whisper by applying a handcrafted signal processing recipe and Voice Conversion (VC) techniques to convert normally phonated speech to whispered speech. We investigate using Gaussian …

Speech SynthesisVoice Conversion

Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition

2025-10-04 · Martin Kocour, Martin Karafiat, Alexander Polok, Dominik Klement 외 arxiv

We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approach leverages a Diarization-Conditioned W…

Speech Recognition

Target Speaker ASR with Whisper

2024-09-14 · Alexander Polok, Dominik Klement, Matthew Wiesner, Sanjeev Khudanpur 외

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speak…

Speech Separation

Extending Whisper with prompt tuning to target-speaker ASR

2023-12-13 · Hao Ma, Zhiyuan Peng, Mingjie Shao, Jing Li 외

Target-speaker automatic speech recognition (ASR) aims to transcribe the desired speech of a target speaker from multi-talker overlapped utterances. Most of the existing target-speaker ASR (TS-ASR) methods involve either…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)parameter-efficient fine-tuningspeech-recognition+2