paper-with-me

홈 › Papers

ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian

2025-11-03 · Nikola Ljubešić, Peter Rupnik, Ivan Porupski, Taja Kuzman Pungeršek arxiv

ParlaSpeech is a collection of spoken parliamentary corpora currently spanning four Slavic languages - Croatian, Czech, Polish and Serbian - all together 6 thousand hours in size. The corpora were built in an automatic fashion from the ParlaMint transcripts and their corresponding metadata, which were aligned to the speech recordings of each corresponding parliament. In this release of the dataset, each of the corpora is significantly enriched with various automatic annotation layers. The textual modality of all four corpora has been enriched with linguistic annotations and sentiment predictions. Similar to that, their spoken modality has been automatically enriched with occurrences of filled pauses, the most frequent disfluency in typical speech. Two out of the four languages have been additionally enriched with detailed word- and grapheme-level alignments, and the automatic annotation of the position of primary stress in multisyllabic words. With these enrichments, the usefulness of the underlying corpora has been drastically increased for downstream research across multiple disciplines, which we showcase through an analysis of acoustic correlates of sentiment. All the corpora are made available for download in JSONL and TextGrid formats, as well as for search through a concordancer.

📄 PDF Abstract BibTeX arXiv:2511.01619

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings

2024-09-23 · Nikola Ljubešić, Peter Rupnik, Danijel Koržinek

Recent significant improvements in speech and language technologies come both from self-supervised approaches over raw language data as well as various types of explicit supervision. To ensure high-quality processing of …

ParlaSpeech-HR - a Freely Available ASR Dataset for Croatian Bootstrapped from the ParlaMint Corpus

2022-06-01 · ParlaCLARIN (LREC) 2022 6 · Nikola Ljubešić, Danijel Koržinek, Peter Rupnik, Ivo-Pavao Jazbec

This paper presents our bootstrapping efforts of producing the first large freely available Croatian automatic speech recognition (ASR) dataset, 1,816 hours in size, obtained from parliamentary transcripts and recordings…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

German Parliamentary Corpus (GerParCor)

2022-04-21 · LREC 2022 6 · Giuseppe Abrami, Mevlüt Bagci, Leon Hammerla, Alexander Mehler

Parliamentary debates represent a large and partly unexploited treasure trove of publicly accessible texts. In the German-speaking area, there is a certain deficit of uniformly accessible and annotated corpora covering a…

Optical Character Recognition (OCR)

FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions

2026-05-26 · Francisco Teixeira, Carlos Carvalho, Mariana Julião, Catarina Botelho 외 arxiv

State-of-the-art performance for Automatic Speech Recognition (ASR) largely depends on the availability of large-scale labeled corpora. This creates a demand for increased data collection efforts, particularly for under-…

Speech Recognition

Parliamentary Discourse Research in Sociology: Literature Review

2022-06-01 · ParlaCLARIN (LREC) 2022 6 · Jure Skubic, Darja Fišer

One of the major sociological research interests has always been the study of political discourse. This literature review gives an overview of the most prominent topics addressed and the most popular methods used by soci…

Sociology