paper-with-me

홈 › Papers

SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding

2025-09-18 · Bingsong Bai, Qihang Lu, Wenbing Yang, Zihan Sun, Yueran Hou, Peilei Jia, Songbai Pu, Ruibo Fu, Yingming Gao, Ya Li, Jun Gao arxiv

Paralinguistic sounds, like laughter and sighs, are crucial for synthesizing more realistic and engaging speech. However, existing methods typically depend on proprietary datasets, while publicly available resources often suffer from incomplete speech, inaccurate or missing timestamps, and limited real-world relevance. To address these problems, we propose an automated framework for generating large-scale paralinguistic data and apply it to construct the SynParaSpeech dataset. The dataset comprises 6 paralinguistic categories with 118.75 hours of data and precise timestamps, all derived from natural conversational speech. Our contributions lie in introducing the first automated method for constructing large-scale paralinguistic datasets and releasing the SynParaSpeech corpus, which advances speech generation through more natural paralinguistic synthesis and enhances speech understanding by improving paralinguistic event detection. The dataset and audio samples are available at https://github.com/ShawnPi233/SynParaSpeech.

📄 PDF Abstract BibTeX arXiv:2509.14946

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Emotion Understanding of Synthesized Speech

2026-03-17 · Yuan Ge, Haishu Zhao, Aokai Hao, Junxiang Zhang 외 arxiv

Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding r…

Speech Emotion RecognitionSpeech Synthesis

DNN-based Speech Synthesis Using Abundant Tags of Spontaneous Speech Corpus

2020-05-01 · LREC 2020 5 · Yuki Yamashita, Tomoki Koriyama, Yuki Saito, Shinnosuke Takamichi 외

In this paper, we investigate the effectiveness of using rich annotations in deep neural network (DNN)-based statistical speech synthesis. DNN-based frameworks typically use linguistic information as input features calle…

Speech Synthesis

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

2025-08-06 · Huan Liao, Qinke Ni, Yuancheng Wang, Yiheng Lu 외 arxiv

Paralinguistic vocalizations-including non-verbal sounds like laughter and breathing, as well as lexicalized interjections such as "uhm" and "oh"-are integral to natural spoken communication. Despite their importance in …

Speech RecognitionSpeech Synthesis

Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations

2025-03-15 · Xue Jiang, Xiulian Peng, Yuan Zhang, Yan Lu

Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, following a semantic-modeling and acoustic-…

Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information

2025-05-21 · Nicholas Sanders, Yuanchao Li, Korin Richmond, Simon King

Quantization in SSL speech models (e.g., HuBERT) improves compression and performance in tasks like language modeling, resynthesis, and text-to-speech but often discards prosodic and paralinguistic information (e.g., emo…

Language ModelingLanguage ModellingQuantizationResynthesis+2