paper-with-me

Papers

LaughNet: synthesizing laughter utterances from waveform silhouettes and a single laughter example

2021-10-11 · Hieu-Thi Luong, Junichi Yamagishi

Emotional and controllable speech synthesis is a topic that has received much attention. However, most studies focused on improving the expressiveness and controllability in the context of linguistic content, even though natural verbal human communication is inseparable from spontaneous non-speech expressions such as laughter, crying, or grunting. We propose a model called LaughNet for synthesizing laughter by using waveform silhouettes as inputs. The motivation is not simply synthesizing new laughter utterances, but testing a novel synthesis-control paradigm that uses an abstract representation of the waveform. We conducted basic listening test experiments, and the results showed that LaughNet can synthesize laughter utterances with moderate quality and retain the characteristics of the training example. More importantly, the generated waveforms have shapes similar to the input silhouettes. For future work, we will test the same method on other types of human nonverbal expressions and integrate it into more elaborated synthesis systems.

📄 PDF Abstract BibTeX arXiv:2110.04946

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Laugh Betrays You? Learning Robust Speaker Representation From Speech Containing Non-Verbal Fragments

2022-10-28 · Yuke Lin, Xiaoyi Qin, Huahua Cui, Zhenyi Zhu 외

The success of automatic speaker verification shows that discriminative speaker representations can be extracted from neutral speech. However, as a kind of non-verbal voice, laughter should also carry speaker information…

Speaker Verification

SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter

2026-05-27 · Lee Jung-Mok, Kim Sung-Bin, Joohyun Chang, Lee Hyun 외 arxiv

Laughter is a complex social signal that conveys communicative intent beyond amusement. While prior work has focused on isolated laughter analysis tasks, a comprehensive understanding of laughter in real-world scenarios …

The AV-LASYN Database : A synchronous corpus of audio and 3D facial marker data for audio-visual laughter synthesis

2014-05-01 · LREC 2014 5 · H{\"u}seyin {\c{C}}akmak, J{\'e}r{\^o}me Urbain, Thierry Dutoit, Jo{\"e}lle Tilmanne

A synchronous database of acoustic and 3D facial marker data was built for audio-visual laughter synthesis. Since the aim is to use this database for HMM-based modeling and synthesis, the amount of collected data from on…

Dimensionality ReductionSpeech Synthesis

MR-RawNet: Speaker verification system with multiple temporal resolutions for variable duration utterances using raw waveforms

2024-06-11 · Seung-bin Kim, Chan-yeong Lim, Jungwoo Heo, Ju-ho Kim 외

In speaker verification systems, the utilization of short utterances presents a persistent challenge, leading to performance degradation primarily due to insufficient phonetic information to characterize the speakers. To…

Speaker Verification

Incremental Disfluency Detection for Spoken Learner English

2022-07-01 · NAACL (BEA) 2022 7 · Lucy Skidmore, Roger Moore

Incremental disfluency detection provides a framework for computing communicative meaning from hesitations, repetitions and false starts commonly found in speech. One application of this area of research is in dialogue-b…