paper-with-me

Papers

Haha-Pod: An Attempt for Laughter-based Non-Verbal Speaker Verification

2023-09-25 · Yuke Lin, Xiaoyi Qin, Ning Jiang, Guoqing Zhao, Ming Li

It is widely acknowledged that discriminative representation for speaker verification can be extracted from verbal speech. However, how much speaker information that non-verbal vocalization carries is still a puzzle. This paper explores speaker verification based on the most ubiquitous form of non-verbal voice, laughter. First, we use a semi-automatic pipeline to collect a new Haha-Pod dataset from open-source podcast media. The dataset contains over 240 speakers' laughter clips with corresponding high-quality verbal speech. Second, we propose a Two-Stage Teacher-Student (2S-TS) framework to minimize the within-speaker embedding distance between verbal and non-verbal (laughter) signals. Considering Haha-Pod as a test set, two trials (S2L-Eval) are designed to verify the speaker's identity through laugh sounds. Experimental results demonstrate that our method can significantly improve the performance of the S2L-Eval test set with only a minor degradation on the VoxCeleb1 test set. The resources for the Haha-Pod dataset can be found at https://github.com/nevermoreLin/HahaPod.

📄 PDF Abstract BibTeX arXiv:2309.14109

Code (1)

nevermorelin/hahapod 공식 구현

Tasks

Speaker Verification

Similar Papers 제목 키워드 기반

Laugh Betrays You? Learning Robust Speaker Representation From Speech Containing Non-Verbal Fragments

2022-10-28 · Yuke Lin, Xiaoyi Qin, Huahua Cui, Zhenyi Zhu 외

The success of automatic speaker verification shows that discriminative speaker representations can be extracted from neutral speech. However, as a kind of non-verbal voice, laughter should also carry speaker information…

Speaker Verification

To laugh or not to laugh? The use of laughter to mark discourse structure

2022-09-01 · SIGDIAL (ACL) 2022 9 · Bogdan Ludusan, Barbara Schuppler

A number of cues, both linguistic and non-linguistic, have been found to mark discourse structure in conversation. This paper investigates the role of laughter, one of the most encountered non-verbal vocalizations in hum…

LaughNet: synthesizing laughter utterances from waveform silhouettes and a single laughter example

2021-10-11 · Hieu-Thi Luong, Junichi Yamagishi

Emotional and controllable speech synthesis is a topic that has received much attention. However, most studies focused on improving the expressiveness and controllability in the context of linguistic content, even though…

Speech Synthesis

A generative framework for conversational laughter: Its 'language model' and laughter sound synthesis

2023-06-06 · Hiroki Mori, Shunya Kimura

As the phonetic and acoustic manifestations of laughter in conversation are highly diverse, laughter synthesis should be capable of accommodating such diversity while maintaining high controllability. This paper proposes…

DiversityLanguage ModelingLanguage Modelling

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech

2025-07-17 · Maksim Borisov, Egor Spirin, Daria Diatlova

Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we introduce NonverbalTTS (NVTTS), a 17-hour …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion ClassificationExpressive Speech Synthesis+5