paper-with-me

홈 › Papers

BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus

2022-07-07 · Josh Meyer, David Ifeoluwa Adelani, Edresson Casanova, Alp Öktem, Daniel Whitenack Julian Weber, Salomon Kabongo, Elizabeth Salesky, Iroro Orife, Colin Leong, Perez Ogayo, Chris Emezue, Jonathan Mukiibi, Salomey Osei, Apelete Agbolo, Victor Akinode, Bernard Opoku, Samuel Olanrewaju, Jesujoba Alabi, Shamsuddeen Muhammad

BibleTTS is a large, high-quality, open speech dataset for ten languages spoken in Sub-Saharan Africa. The corpus contains up to 86 hours of aligned, studio quality 48kHz single speaker recordings per language, enabling the development of high-quality text-to-speech models. The ten languages represented are: Akuapem Twi, Asante Twi, Chichewa, Ewe, Hausa, Kikuyu, Lingala, Luganda, Luo, and Yoruba. This corpus is a derivative work of Bible recordings made and released by the Open.Bible project from Biblica. We have aligned, cleaned, and filtered the original recordings, and additionally hand-checked a subset of the alignments for each language. We present results for text-to-speech models with Coqui TTS. The data is released under a commercial-friendly CC-BY-SA license.

📄 PDF Abstract BibTeX arXiv:2207.03546

Code (1)

alpoktem/bible2speechdb 공식 구현

Tasks

text-to-speechText to SpeechVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

2026-06-08 · David Guzmán, Luel Hagos Beyene, Jesujoba Oluwadara Alabi, Yejin Jeon 외 arxiv

Recent advances in neural text-to-speech (TTS) and multilingual speech generation have substantially improved synthetic speech quality, yet these gains remain unevenly distributed across the world's languages. Existing m…

Speech Synthesis

SUTRA: Scalable Multilingual Language Model Architecture

2024-05-07 · Abhijit Bendale, Michael Sapienza, Steven Ripplinger, Simon Gibbs 외

In this paper, we introduce SUTRA, multilingual Large Language Model architecture capable of understanding, reasoning, and generating text in over 50 languages. SUTRA's design uniquely decouples core conceptual understan…

Computational EfficiencyHallucinationLanguage ModelingLanguage Modelling+4

KInIT at SemEval-2024 Task 8: Fine-tuned LLMs for Multilingual Machine-Generated Text Detection

2024-02-21 · Michal Spiegel, Dominik Macko

SemEval-2024 Task 8 is focused on multigenerator, multidomain, and multilingual black-box machine-generated text detection. Such a detection is important for preventing a potential misuse of large language models (LLMs),…

Language Identificationparameter-efficient fine-tuningtext-classificationText Classification+1

Breaking Language Barriers in Visual Language Models via Multilingual Textual Regularization

2025-03-28 · Iñigo Pikabea, Iñaki Lacunza, Oriol Pareras, Carlos Escolano 외

Rapid advancements in Visual Language Models (VLMs) have transformed multimodal understanding but are often constrained by generating English responses regardless of the input language. This phenomenon has been termed as…

On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons

2024-04-03 · Takeshi Kojima, Itsuki Okimura, Yusuke Iwasawa, Hitomi Yanaka 외

Current decoder-based pre-trained language models (PLMs) successfully demonstrate multilingual capabilities. However, it is unclear how these models handle multilingualism. We analyze the neuron-level internal behavior o…

DecoderText Generation