paper-with-me

홈 › Papers

Target Speaker ASR with Whisper

2024-09-14 · Alexander Polok, Dominik Klement, Matthew Wiesner, Sanjeev Khudanpur, Jan Černocký, Lukáš Burget

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to condition on frame-level diarization outputs than to learn the space of all speaker embeddings. We find that adding even a single bias term per diarization output type before the first transformer block can transform single-speaker ASR models into target-speaker ASR models. Our approach also supports speaker-attributed ASR by sequentially generating transcripts for each speaker in a diarization output. This simplified method outperforms baseline speech separation and diarization cascade by 12.9 % absolute ORC-WER on the NOTSOFAR-1 dataset.

📄 PDF Abstract BibTeX arXiv:2409.09543

Code (1)

BUTSpeechFIT/TS-ASR-Whisper 공식 구현 pytorch

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR

2024-12-07 · Pengcheng Guo, Xuankai Chang, Hang Lv, Shinji Watanabe 외

Benefiting from massive and diverse data sources, speech foundation models exhibit strong generalization and knowledge transfer capabilities to a wide range of downstream tasks. However, a limitation arises from their ex…

Automatic Speech RecognitionData Augmentationspeech-recognitionSpeech Recognition+1

DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition

2024-12-30 · Alexander Polok, Dominik Klement, Martin Kocour, Jiangyu Han 외

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a significant challenge, particularly when systems conditioned on speaker embeddings fail to generalize to unseen speakers. In t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2

Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition

2025-10-04 · Martin Kocour, Martin Karafiat, Alexander Polok, Dominik Klement 외 arxiv

We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approach leverages a Diarization-Conditioned W…

Speech Recognition

Voice Conversion for Whispered Speech Synthesis

2019-12-11 · Marius Cotescu, Thomas Drugman, Goeric Huybrechts, Jaime Lorenzo-Trueba 외

We present an approach to synthesize whisper by applying a handcrafted signal processing recipe and Voice Conversion (VC) techniques to convert normally phonated speech to whispered speech. We investigate using Gaussian …

Speech SynthesisVoice Conversion

Extending Whisper with prompt tuning to target-speaker ASR

2023-12-13 · Hao Ma, Zhiyuan Peng, Mingjie Shao, Jing Li 외

Target-speaker automatic speech recognition (ASR) aims to transcribe the desired speech of a target speaker from multi-talker overlapped utterances. Most of the existing target-speaker ASR (TS-ASR) methods involve either…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)parameter-efficient fine-tuningspeech-recognition+2