paper-with-me

홈 › Papers

NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion

2025-10-31 · Zongyang Du, Shreeram Suresh Chandra, Ismail Rasim Ulgen, Aurosweta Mahapatra, Ali N. Salman, Carlos Busso, Berrak Sisman arxiv

Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted, limited in scale, and fail to capture the expressive richness of real-life communication. With the rise of large neural networks, several large-scale speech corpora have emerged and been widely adopted across various speech processing tasks. However, the field of voice conversion (VC) still lacks large-scale, expressive, and real-life speech resources suitable for modeling natural prosody and emotion. To fill this gap, we release NaturalVoices (NV), the first large-scale spontaneous podcast dataset specifically designed for emotion-aware voice conversion. It comprises 5,049 hours of spontaneous podcast recordings with automatic annotations for emotion (categorical and attribute-based), speech quality, transcripts, speaker identity, and sound events. The dataset captures expressive emotional variation across thousands of speakers, diverse topics, and natural speaking styles. We also provide an open-source pipeline with modular annotation tools and flexible filtering, enabling researchers to construct customized subsets for a wide range of VC tasks. Experiments demonstrate that NaturalVoices supports the development of robust and generalizable VC models capable of producing natural, expressive speech, while revealing limitations of current architectures when applied to large-scale spontaneous data. These results suggest that NaturalVoices is both a valuable resource and a challenging benchmark for advancing the field of voice conversion. Dataset is available at: https://huggingface.co/JHU-SmileLab

📄 PDF Abstract BibTeX arXiv:2511.00256

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline

2024-06-06 · Ali N. Salman, Zongyang Du, Shreeram Suresh Chandra, Ismail Rasim Ulgen 외

Voice conversion (VC) research traditionally depends on scripted or acted speech, which lacks the natural spontaneity of real-life conversations. While natural speech data is limited for VC, our study focuses on filling …

Voice Conversion

MoonCast: High-Quality Zero-Shot Podcast Generation

2025-03-18 · Zeqian Ju, Dongchao Yang, Jianwei Yu, Kai Shen 외

Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers. However, these systems still face challenges when extending their capabilitie…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

SwissGPC v1.0 -- The Swiss German Podcasts Corpus

2025-09-24 · Samuel Stucki, Mark Cieliebak, Jan Deriu arxiv

We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to ta…

Classification of Spontaneous and Scripted Speech for Multilingual Audio

2024-12-16 · Shahar Elisha, Andrew McDowell, Mariano Beguerisse-Díaz, Emmanouil Benetos

Distinguishing scripted from spontaneous speech is an essential tool for better understanding how speech styles influence speech processing research. It can also improve recommendation systems and discovery experiences f…

Recommendation Systems

Podcasts as a Medium for Participation in Collective Action: A Case Study of Black Lives Matter

2025-09-16 · Theodora Moldovan, Arianna Pera, Davide Vega, Luca Maria Aiello arxiv

We study how participation in collective action is articulated in podcast discussions, using the Black Lives Matter (BLM) movement as a case study. While research on collective action discourse has primarily focused on t…