paper-with-me

Papers

Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation

2025-09-18 · Miseul Kim, Soo Jin Park, Kyungguen Byun, Hyeon-Kyeong Shin, Sunkuk Moon, Shuhua Zhang, Erik Visser arxiv

Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different individuals, for example, when one raises their voice or speaks faster during conversation. To address this, we propose a style-controllable speech generation model that augments speech across diverse styles while preserving the target speaker's identity. The proposed system starts with diarized segments from a conventional diarizer. For each diarized segment, it generates augmented speech samples enriched with phonetic and stylistic diversity. And then, speaker embeddings from both the original and generated audio are blended to enhance the system's robustness in grouping segments with high intrinsic intra-speaker variability. We validate our approach on a simulated emotional speech dataset and the truncated AMI dataset, demonstrating significant improvements, with error rate reductions of 49% and 35% on each dataset, respectively.

📄 PDF Abstract BibTeX arXiv:2509.14632

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Diarization

Similar Papers 제목 키워드 기반

Constrained speaker diarization of TV series based on visual patterns

2018-12-18 · Xavier Bost, Georges Linares

Speaker diarization, usually denoted as the ''who spoke when'' task, turns out to be particularly challenging when applied to fictional films, where many characters talk in various acoustic conditions (background music, …

Clusteringspeaker-diarizationSpeaker Diarization

Audiovisual speaker diarization of TV series

2018-12-18 · Xavier Bost, Georges Linarès, Serigne Gueye

Speaker diarization may be difficult to achieve when applied to narrative films, where speakers usually talk in adverse acoustic conditions: background music, sound effects, wide variations in intonation may hide the int…

speaker-diarizationSpeaker Diarization

The Third DIHARD Diarization Challenge

2020-12-02 · Neville Ryant, Prachi Singh, Venkat Krishnamohan, Rajat Varma 외

DIHARD III was the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variability in recording equipment, noise conditions, and conversational domain. Speaker…

speaker-diarizationSpeaker Diarizationvalid

D{é}tection de locuteurs dans les s{é}ries TV

2018-12-18 · Xavier Bost, Georges Linares

Speaker diarization of audio streams turns out to be particularly challenging when applied to fictional films, where many characters talk in various acoustic conditions (background music, sound effects, variations in int…

Clusteringspeaker-diarizationSpeaker Diarization

Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization

2026-05-06 · Mohammed Aman Bhuiyan, Md Sazzad Hossain Adib, Samiul Basir Bhuiyan, Amit Chakraborty 외 arxiv

Automatic Speech Recognition (ASR) and speaker diarization in Bangla remain challenging due to long form recordings, diverse acoustic conditions, and significant speaker variability. This work addresses these two core ta…

Spoken Language UnderstandingSpeaker DiarizationSpeech RecognitionData Augmentation