paper-with-me

홈 › Papers

OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics

2025-09-04 · Wei Chu, Yuanzhe Dong, Ke Tan, Dong Han, Xavier Menendez-Pidal, Ruchao Fan, Chenfeng Miao, Chanwoo Kim, Bhiksha Raj, Rita Singh arxiv

OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podcasts, talk shows, teleconferences, and other conversations. Speaker names, turns, and transcripts are human-sourced and refined by a proprietary pipeline, while additional information such as timestamps and confidence scores is derived from the pipeline. The IV denotes its position as Tier IV in the Olewave dataset series. In addition, we have open-sourced a subset, OleSpeech-IV-2025-EN-AR-100, for non-commercial research use.

📄 PDF Abstract BibTeX arXiv:2509.04702

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings

2020-04-20 · Shinji Watanabe, Michael Mandel, Jon Barker, Emmanuel Vincent 외

Following the success of the 1st, 2nd, 3rd, 4th and 5th CHiME challenges we organize the 6th CHiME Speech Separation and Recognition Challenge (CHiME-6). The new challenge revisits the previous CHiME-5 challenge and furt…

speaker-diarizationSpeaker DiarizationSpeech Enhancementspeech-recognition+2

Multilingual Multiaccented Multispeaker TTS with RADTTS

2023-01-24 · Rohan Badlani, Rafael Valle, Kevin J. Shih, João Felipe Santos 외

We work to create a multilingual speech synthesis system which can generate speech with the proper accent while retaining the characteristics of an individual voice. This is challenging to do because it is expensive to o…

Speech Synthesis

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

2026-07-25 · Sahil Deepak Gawande, Mayank Singh arxiv

Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between Engl…

Dialogue Generation

Speaker verification-derived loss and data augmentation for DNN-based multispeaker speech synthesis

2021-06-03 · Beata Lorincz, Adriana Stan, Mircea Giurgiu

Building multispeaker neural network-based text-to-speech synthesis systems commonly relies on the availability of large amounts of high quality recordings from each speaker and conditioning the training process on the s…

Data AugmentationSpeaker VerificationSpeech Synthesistext-to-speech+2

Low-Resource Multilingual and Zero-Shot Multispeaker TTS

2022-10-21 · Florian Lux, Julia Koch, Ngoc Thang Vu

While neural methods for text-to-speech (TTS) have shown great advances in modeling multiple speakers, even in zero-shot settings, the amount of data needed for those approaches is generally not feasible for the vast maj…

Meta-Learningtext-to-speechText to SpeechVoice Cloning