paper-with-me

Papers

Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

2026-06-19 · Tzu-Chieh Wei, Yi-Cheng Lin, Huang-Cheng Chou, Kuan-Yu Chen, Hsin-Yen Sung, Shrikanth Narayanan, Hung-yi Lee arxiv

As expressive text-to-speech (TTS) and voice conversion (VC) systems increasingly generate non-verbal vocalizations (NVVs) to enhance naturalness, reliable speaker verification (SV) becomes essential to objectively assess identity consistency across both verbal and non-verbal segments. Yet current SV systems generalize poorly to NVVs, and fine-tuning on NVV data causes catastrophic forgetting of speech performance. We present the first systematic study across 10 NVV types and propose a framework combining frozen Data2Vec self-supervised features with ECAPA-TDNN, enhanced by a Mixture of Experts (MoE) module with learned domain-aware routing. A conditional distillation loss on speech inputs via a pretrained teacher retains speech-to-speech accuracy, while a contrastive loss bridges the speech-NVV domain gap. Our method reduces speech-NVV EER from 38.93% to 22.66% over a pretrained baseline, and improves speech EER from 13.17% to 9.24% via distillation.

📄 PDF Abstract BibTeX arXiv:2606.21215

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationVoice Conversion

Similar Papers 제목 키워드 기반

Proceedings of the ICML 2022 Expressive Vocalizations Workshop and Competition: Recognizing, Generating, and Personalizing Vocal Bursts

2022-07-14 · Alice Baird, Panagiotis Tzirakis, Gauthier Gidel, Marco Jiralerspong 외

This is the Proceedings of the ICML Expressive Vocalization (ExVo) Competition. The ExVo competition focuses on understanding and generating vocal bursts: laughs, gasps, cries, and other non-verbal vocalizations that are…

Few-Shot Learning

The ICML 2022 Expressive Vocalizations Workshop and Competition: Recognizing, Generating, and Personalizing Vocal Bursts

2022-05-03 · Alice Baird, Panagiotis Tzirakis, Gauthier Gidel, Marco Jiralerspong 외

The ICML Expressive Vocalization (ExVo) Competition is focused on understanding and generating vocal bursts: laughs, gasps, cries, and other non-verbal vocalizations that are central to emotional expression and communica…

Few-Shot Learning

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech

2025-07-17 · Maksim Borisov, Egor Spirin, Daria Diatlova

Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we introduce NonverbalTTS (NVTTS), a 17-hour …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion ClassificationExpressive Speech Synthesis+5

Textless Speech Emotion Conversion using Decomposed & Discrete Representations

2021-11-14 · arXiv 2021 11 · Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov 외

Speech emotion conversion is the task of modifying the perceived emotion of a speech utterance while preserving the lexical content and speaker identity. In this study, we cast the problem of emotion conversion as a spok…

Haha-Pod: An Attempt for Laughter-based Non-Verbal Speaker Verification

2023-09-25 · Yuke Lin, Xiaoyi Qin, Ning Jiang, Guoqing Zhao 외

It is widely acknowledged that discriminative representation for speaker verification can be extracted from verbal speech. However, how much speaker information that non-verbal vocalization carries is still a puzzle. Thi…

Speaker Verification