paper-with-me

Papers

SpeakGer: A meta-data enriched speech corpus of German state and federal parliaments

2024-10-23 · Kai-Robin Lange, Carsten Jentsch

The application of natural language processing on political texts as well as speeches has become increasingly relevant in political sciences due to the ability to analyze large text corpora which cannot be read by a single person. But such text corpora often lack critical meta information, detailing for instance the party, age or constituency of the speaker, that can be used to provide an analysis tailored to more fine-grained research questions. To enable researchers to answer such questions with quantitative approaches such as natural language processing, we provide the SpeakGer data set, consisting of German parliament debates from all 16 federal states of Germany as well as the German Bundestag from 1947-2023, split into a total of 10,806,105 speeches. This data set includes rich meta data in form of information on both reactions from the audience towards the speech as well as information about the speaker's party, their age, their constituency and their party's political alignment, which enables a deeper analysis. We further provide three exploratory analyses, detailing topic shares of different parties throughout time, a descriptive analysis of the development of the age of an average speaker as well as a sentiment analysis of speeches of different parties with regards to the COVID-19 pandemic.

📄 PDF Abstract BibTeX arXiv:2410.17886

Code (1)

K-RLange/SpeakGer 공식 구현

Tasks

DescriptiveSentiment Analysis

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets

2025-11-17 · Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik, Tekla Etelka Gráczi 외 arxiv

The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepresented due to limited spontaneous and c…

Speaker DiarizationSpeech Recognition

Two Centuries of Sexism in British Parliament: A Computational Analysis of Women's Representation in the Hansard Corpus

2026-08-31 · Mohammad Omar Khursheed, Mandira Sawkar, Ashiqur R. KhudaBukhsh arxiv

The language a legislature uses to debate women's rights, even in favour of them, encodes systematic patterns of sexism that persist across two centuries. In this work, we analyse 6,531 speeches over 200 years of UK parl…

Enhanced CORILGA: Introducing the Automatic Phonetic Alignment Tool for Continuous Speech

2016-05-01 · LREC 2016 5 · Roberto Seara, Marta Martinez, Roc{\'\i}o Varela, Carmen Garc{\'\i}a Mateo 외

The {``}Corpus Oral Informatizado da Lingua Galega (CORILGA){''} project aims at building a corpus of oral language for Galician, primarily designed to study the linguistic variation and change. This project is currently…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+1

Modeling Turn-Taking with Semantically Informed Gestures

2025-10-22 · Varsha Suresh, M. Hamza Mughal, Christian Theobalt, Vera Demberg arxiv

In conversation, humans use multimodal cues, such as speech, gestures, and gaze, to manage turn-taking. While linguistic and acoustic features are informative, gestures provide complementary cues for modeling these trans…

The Esethu Framework: Reimagining Sustainable Dataset Governance and Curation for Low-Resource Languages

2025-02-21 · Jenalea Rajab, Anuoluwapo Aremu, Everlyn Asiko Chimoto, Dale Dunbar 외

This paper presents the Esethu Framework, a sustainable data curation framework specifically designed to empower local communities and ensure equitable benefit-sharing from their linguistic resources. This framework is s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition