paper-with-me

홈 › Papers

SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models

2025-08-08 · Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng, Hui Wang, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li arxiv

The Speaker Diarization and Recognition (SDR) task aims to predict "who spoke when and what" within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers.

📄 PDF Abstract BibTeX arXiv:2508.06372

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker DiarizationSpeech Recognition

Similar Papers 제목 키워드 기반

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

2025-07-25 · Hsuan-Yu Wang, Pei-Ying Lee, Berlin Chen arxiv

In this paper, we investigate the impact of incorporating timestamp-based alignment between Automatic Speech Recognition (ASR) transcripts and Speaker Diarization (SD) outputs on Speech Emotion Recognition (SER) accuracy…

Multimodal Emotion RecognitionSpeech Emotion RecognitionSpeaker DiarizationSpeech Recognition

3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization

2024-03-29 · Yafeng Chen, Siqi Zheng, Hui Wang, Luyao Cheng 외

We introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit ade…

Self-Supervised Learningspeaker-diarizationSpeaker DiarizationSpeaker Recognition+1

Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition

2023-12-18 · Peng Shen, Xugang Lu, Hisashi Kawai

Multi-talker overlapped speech recognition remains a significant challenge, requiring not only speech recognition but also speaker diarization tasks to be addressed. In this paper, to better address these tasks, we first…

speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition

Multimodal Speaker Segmentation and Diarization using Lexical and Acoustic Cues via Sequence to Sequence Neural Networks

2018-05-28

While there has been substantial amount of work in speaker diarization recently, there are few efforts in jointly employing lexical and acoustic information for speaker segmentation. Towards that, we investigate a speake…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2

Speaker Diarization with Lexical Information

2020-04-13 · Tae Jin Park, Kyu J. Han, Jing Huang, Xiaodong He 외

This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeaker-diarization+3