paper-with-me

홈 › Papers

Tagarela - A Portuguese speech dataset from podcasts

2026-03-16 · Frederico Santos de Oliveira, Lucas Rafael Stefanel Gris, Alef Iury Siqueira Ferreira, Augusto Seben da Rosa, Alexandre Costa Ferro Filho, Edresson Casanova, Christopher Dane Shulby, Rafael Teixeira Sousa, Diogo Fernandes Costa Silva, Anderson da Silva Soares, Arlindo Rodrigues Galvão Filho arxiv

Despite significant advances in speech processing, Portuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over 8,972 hours of podcast audio, specifically curated for training automatic speech recognition (ASR) and text-to-speech (TTS) models. Notably, its scale rivals English's GigaSpeech (10kh), enabling state-of-the-art Portuguese models. To ensure data quality, the corpus was subjected to an audio pre-processing pipeline and subsequently transcribed using a mixed strategy: we applied ASR models that were previously trained on high-fidelity transcriptions generated by proprietary APIs, ensuring a high level of initial accuracy. Finally, to validate the effectiveness of this new resource, we present ASR and TTS models trained exclusively on our dataset and evaluate their performance, demonstrating its potential to drive the development of more robust and natural speech technologies for Portuguese. The dataset is released publicly, available at https://freds0.github.io/TAGARELA/, to foster the development of robust speech technologies.

📄 PDF Abstract BibTeX arXiv:2603.15326

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Cem Mil Podcasts: A Spoken Portuguese Document Corpus For Multi-modal, Multi-lingual and Multi-Dialect Information Access Research

2022-09-23 · Ekaterina Garmash, Edgar Tanaka, Ann Clifton, Joana Correia 외

In this paper we describe the Portuguese-language podcast dataset we have released for academic research purposes. We give an overview of how the data was sampled, descriptive statistics over the collection, as well as i…

DescriptiveGenre classification

SEP-28k: A Dataset for Stuttering Event Detection From Podcasts With People Who Stutter

2021-02-24 · Colin Lea, Vikramjit Mitra, Aparna Joshi, Sachin Kajarekar 외

The ability to automatically detect stuttering events in speech could help speech pathologists track an individual's fluency over time or help improve speech recognition systems for people with atypical speech patterns. …

Event Detectionspeech-recognitionSpeech Recognition

100,000 Podcasts: A Spoken English Document Corpus

2020-12-01 · COLING 2020 8 · Ann Clifton, Sravana Reddy, Yongze Yu, Aasish Pappu 외

Podcasts are a large and growing repository of spoken audio. As an audio format, podcasts are more varied in style and production type than broadcast news, contain more genres than typically studied in video data, and ar…

3D Facial Landmark LocalizationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Facial Expression Recognition (FER)+5

SwissGPC v1.0 -- The Swiss German Podcasts Corpus

2025-09-24 · Samuel Stucki, Mark Cieliebak, Jan Deriu arxiv

We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to ta…

Transcribe, Align and Segment: Creating speech datasets for low-resource languages

2024-06-18 · Taras Sereda

In this work, we showcase a cost-effective method for generating training data for speech processing tasks. First, we transcribe unlabeled speech using a state-of-the-art Automatic Speech Recognition (ASR) model. Next, w…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition