paper-with-me

홈 › Papers

voc2vec: A Foundation Model for Non-Verbal Vocalization

2025-02-22 · Alkis Koudounas, Moreno La Quatra, Marco Sabato Siniscalchi, Elena Baralis

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various real-world applications. Audio foundation models well handle non-speech data but also fail to capture the nuanced features of non-verbal human sounds. In this work, we aim to overcome the above shortcoming and propose a novel foundation model, termed voc2vec, specifically designed for non-verbal human data leveraging exclusively open-source non-verbal audio datasets. We employ a collection of 10 datasets covering around 125 hours of non-verbal audio. Experimental results prove that voc2vec is effective in non-verbal vocalization classification, and it outperforms conventional speech and audio foundation models. Moreover, voc2vec consistently outperforms strong baselines, namely OpenSmile and emotion2vec, on six different benchmark datasets. To the best of the authors' knowledge, voc2vec is the first universal representation model for vocalization tasks.

📄 PDF Abstract BibTeX arXiv:2502.16298

Code (1)

koudounasalkis/voc2vec 공식 구현 pytorch

Tasks

model

Similar Papers 제목 키워드 기반

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

2025-09-19 · Jialong Mai, Jinxin Ji, Xiaofen Xing, Chen Yang 외 arxiv

Mainstream Automatic Speech Recognition (ASR) systems excel at transcribing lexical content, but largely fail to recognize nonverbal vocalizations (NVs) embedded in speech, such as sighs, laughs, and coughs. This capabil…

Speech Recognition

Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

2026-06-19 · Tzu-Chieh Wei, Yi-Cheng Lin, Huang-Cheng Chou, Kuan-Yu Chen 외 arxiv

As expressive text-to-speech (TTS) and voice conversion (VC) systems increasingly generate non-verbal vocalizations (NVVs) to enhance naturalness, reliable speaker verification (SV) becomes essential to objectively asses…

Speaker VerificationVoice Conversion

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech

2025-07-17 · Maksim Borisov, Egor Spirin, Daria Diatlova

Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we introduce NonverbalTTS (NVTTS), a 17-hour …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion ClassificationExpressive Speech Synthesis+5

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

2025-08-06 · Huan Liao, Qinke Ni, Yuancheng Wang, Yiheng Lu 외 arxiv

Paralinguistic vocalizations-including non-verbal sounds like laughter and breathing, as well as lexicalized interjections such as "uhm" and "oh"-are integral to natural spoken communication. Despite their importance in …

Speech RecognitionSpeech Synthesis

Proceedings of the ICML 2022 Expressive Vocalizations Workshop and Competition: Recognizing, Generating, and Personalizing Vocal Bursts

2022-07-14 · Alice Baird, Panagiotis Tzirakis, Gauthier Gidel, Marco Jiralerspong 외

This is the Proceedings of the ICML Expressive Vocalization (ExVo) Competition. The ExVo competition focuses on understanding and generating vocal bursts: laughs, gasps, cries, and other non-verbal vocalizations that are…

Few-Shot Learning