paper-with-me

홈 › Papers

TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding

2026-01-11 · Mingyue Huo, Yiwen Shao, Yuheng Zhang arxiv

We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via Serialized Output Training (SOT) to learn turn-taking dynamics; and (2) an interleaved time anchor mechanism that not only supports fine-grained timestamp prediction but also acts as a synchronization signal between semantic understanding and speaker tracking. Compared to previous works that primarily focus on speaker-attributed ASR or implicit diarization, TagSpeech addresses the challenge of fine-grained speaker-content alignment and explicitly models "who spoke what and when" in an end-to-end manner. Experiments on AMI and AliMeeting benchmarks demonstrate that our method achieves consistent improvements in Diarization Error Rate (DER) over strong end-to-end baselines, including Qwen-Omni and Gemini, particularly in handling complex speech overlaps. Moreover, TagSpeech employs a parameter-efficient training paradigm in which the LLM backbone is frozen and only lightweight projectors are trained, resulting in strong performance with low computational cost.

📄 PDF Abstract BibTeX arXiv:2601.06896

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SDBench: A Comprehensive Benchmark Suite for Speaker Diarization

2025-07-22 · Eduardo Pacheco, Atila Orhon, Berkin Durmus, Blaise Munyampirwa 외 arxiv

Even state-of-the-art speaker diarization systems exhibit high variance in error rates across different datasets, representing numerous use cases and domains. Furthermore, comparing across systems requires careful applic…

Speaker Diarization

Designing an Effective Metric Learning Pipeline for Speaker Diarization

2018-11-01 · Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan, Huan Song, Andreas Spanias

State-of-the-art speaker diarization systems utilize knowledge from external data, in the form of a pre-trained distance metric, to effectively determine relative speaker identities to unseen data. However, much of recen…

Metric Learningspeaker-diarizationSpeaker Diarization

DiarizationLM: Speaker Diarization Post-Processing with Large Language Models

2024-01-07 · Quan Wang, Yiling Huang, Guanlong Zhao, Evan Clark 외

In this paper, we introduce DiarizationLM, a framework to leverage large language models (LLM) to post-process the outputs from a speaker diarization system. Various goals can be achieved with the proposed framework, suc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2

Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization

2026-05-06 · Mohammed Aman Bhuiyan, Md Sazzad Hossain Adib, Samiul Basir Bhuiyan, Amit Chakraborty 외 arxiv

Automatic Speech Recognition (ASR) and speaker diarization in Bangla remain challenging due to long form recordings, diverse acoustic conditions, and significant speaker variability. This work addresses these two core ta…

Spoken Language UnderstandingSpeaker DiarizationSpeech RecognitionData Augmentation

TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization

2023-03-08 · JiaMing Wang, Zhihao Du, Shiliang Zhang

Recently, end-to-end neural diarization (EEND) is introduced and achieves promising results in speaker-overlapped scenarios. In EEND, speaker diarization is formulated as a multi-label prediction problem, where speaker a…

speaker-diarizationSpeaker DiarizationVocal Bursts Valence Prediction