paper-with-me

Papers

The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings

2024-09-23 · Nikola Ljubešić, Peter Rupnik, Danijel Koržinek

Recent significant improvements in speech and language technologies come both from self-supervised approaches over raw language data as well as various types of explicit supervision. To ensure high-quality processing of spoken data, the most useful type of explicit supervision is still the alignment between the speech signal and its corresponding text transcript, which is a data type that is not available for many languages. In this paper, we present our approach to building large and open speech-and-text-aligned datasets of less-resourced languages based on transcripts of parliamentary proceedings and their recordings. Our starting point are the ParlaMint comparable corpora of transcripts of parliamentary proceedings of 26 national European parliaments. In the pilot run on expanding the ParlaMint corpora with aligned publicly available recordings, we focus on three Slavic languages, namely Croatian, Polish, and Serbian. The main challenge of our approach is the lack of any global alignment between the ParlaMint texts and the available recordings, as well as the sometimes varying data order in each of the modalities, which requires a novel approach in aligning long sequences of text and audio in a large search space. The results of this pilot run are three high-quality datasets that span more than 5,000 hours of speech and accompanying text transcripts. Although these datasets already make a huge difference in the availability of spoken and textual data for the three languages, we want to emphasize the potential of the presented approach in building similar datasets for many more languages.

📄 PDF Abstract BibTeX arXiv:2409.15397

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian

2025-11-03 · Nikola Ljubešić, Peter Rupnik, Ivan Porupski, Taja Kuzman Pungeršek arxiv

ParlaSpeech is a collection of spoken parliamentary corpora currently spanning four Slavic languages - Croatian, Czech, Polish and Serbian - all together 6 thousand hours in size. The corpora were built in an automatic f…

ASR-Generated Text for Language Model Pre-training Applied to Speech Tasks

2022-07-05 · Valentin Pelloin, Franck Dary, Nicolas Herve, Benoit Favre 외

We aim at improving spoken language modeling (LM) using very large amount of automatically transcribed speech. We leverage the INA (French National Audiovisual Institute) collection and obtain 19GB of text after applying…

Language ModelingLanguage ModellingSpoken Language Understanding

Using ASR-Generated Text for Spoken Language Modeling

2022-05-01 · BigScience (ACL) 2022 5 · Nicolas Hervé, Valentin Pelloin, Benoit Favre, Franck Dary 외

This papers aims at improving spoken language modeling (LM) using very large amount of automatically transcribed speech. We leverage the INA (French National Audiovisual Institute) collection and obtain 19GB of text afte…

Language ModelingLanguage Modelling

ParlaSpeech-HR - a Freely Available ASR Dataset for Croatian Bootstrapped from the ParlaMint Corpus

2022-06-01 · ParlaCLARIN (LREC) 2022 6 · Nikola Ljubešić, Danijel Koržinek, Peter Rupnik, Ivo-Pavao Jazbec

This paper presents our bootstrapping efforts of producing the first large freely available Croatian automatic speech recognition (ASR) dataset, 1,816 hours in size, obtained from parliamentary transcripts and recordings…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition

2022-03-29 · Junrui Ni, Liming Wang, Heting Gao, Kaizhi Qian 외

An unsupervised text-to-speech synthesis (TTS) system learns to generate speech waveforms corresponding to any written sentence in a language by observing: 1) a collection of untranscribed speech waveforms in that langua…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+5