paper-with-me

홈 › Papers

Multi-aspect Multilingual and Cross-lingual Parliamentary Speech Analysis

2022-07-03 · Kristian Miok, Encarnacion Hidalgo-Tenorio, Petya Osenova, Miguel-Angel Benitez-Castro, Marko Robnik-Sikonja

Parliamentary and legislative debate transcripts provide informative insight into elected politicians' opinions, positions, and policy preferences. They are interesting for political and social sciences as well as linguistics and natural language processing (NLP) research. While existing research studied individual parliaments, we apply advanced NLP methods to a joint and comparative analysis of six national parliaments (Bulgarian, Czech, French, Slovene, Spanish, and United Kingdom) between 2017 and 2020. We analyze emotions and sentiment in the transcripts from the ParlaMint dataset collection and assess if the age, gender, and political orientation of speakers can be detected from their speeches. The results show some commonalities and many surprising differences among the analyzed countries.

📄 PDF Abstract BibTeX arXiv:2207.01054

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings

2023-09-18 · Michal Mochtak, Peter Rupnik, Nikola Ljubešić

The paper presents a new training dataset of sentences in 7 languages, manually annotated for sentiment, which are used in a series of experiments focused on training a robust sentiment identifier for parliamentary proce…

Decision MakingLanguage ModelingLanguage ModellingSentiment Analysis

Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification

2026-02-18 · Taja Kuzman Pungeršek, Peter Rupnik, Daniela Širinić, Nikola Ljubešić arxiv

This paper introduces ParlaCAP, a large-scale dataset for analyzing parliamentary agenda setting across Europe, and proposes a cost-effective method for building domain-specific policy topic classifiers. Applying the Com…

Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

2026-06-18 · Himath Dhanapala, Haren Daishika, Himandhi Kuruppu, Sithija Seneviratne 외 arxiv

Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex…

WorldSpeech: A Multilingual Speech Corpus from Around the World

2026-05-09 · Antonis Asonitis, Luca A. Lanzendörfer, Frédéric Berdoz, Roger Wattenhofer arxiv

Automatic speech recognition (ASR) performs well for high-resource languages with abundant paired audio-transcript data, but its accuracy degrades sharply for most languages due to limited publicly available aligned data…

Speech Recognition

Sri Lanka Document Datasets: A Large-Scale, Multilingual Resource for Law, News, and Policy

2025-10-05 · Nuwan I. Senaratna arxiv

We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comp…