AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI
Political discourse datasets are important for gaining political insights, analyzing communication strategies or social science phenomena. Although numerous political discourse corpora exist, comprehensive, high-quality, annotated datasets are scarce. This is largely due to the substantial manual effort, multidisciplinarity, and expertise required for the nuanced annotation of rhetorical strategies and ideological contexts. In this paper, we present AgoraSpeech, a meticulously curated, high-quality dataset of 171 political speeches from six parties during the Greek national elections in 2023. The dataset includes annotations (per paragraph) for six natural language processing (NLP) tasks: text classification, topic identification, sentiment analysis, named entity recognition, polarization and populism detection. A two-step annotation was employed, starting with ChatGPT-generated annotations and followed by exhaustive human-in-the-loop validation. The dataset was initially used in a case study to provide insights during the pre-election period. However, it has general applicability by serving as a rich source of information for political and social scientists, journalists, or data scientists, while it can be used for benchmarking and fine-tuning NLP and large language models (LLMs).
Code (0)
등록된 구현이 없습니다.
Tasks
Benchmarkingnamed-entity-recognitionNamed Entity RecognitionSentiment Analysistext-classificationText ClassificationSimilar Papers 제목 키워드 기반
Analyzing the Impact of Fake News on the Anticipated Outcome of the 2024 Election Ahead of Time
Despite increasing awareness and research around fake news, there is still a significant need for datasets that specifically target racial slurs and biases within North American political speeches. This is particulary im…
ArticlesBenchmarkingLanguage ModelingLanguage Modelling+1Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles
Detecting media bias is crucial, specifically in the South Asian region. Despite this, annotated datasets and computational studies for Bangla political bias research remain scarce. Crucially because, political stance de…
Stance DetectionArPoMeme: An Annotated Arabic Multimodal Dataset for Political Ideology and Polarization
Memes have become a prominent medium of political communication in the Arab world, reflecting how humor, imagery, and text interact to express ideological and cultural positions. Despite the centrality of memes to online…
Does Twitter know your political views? POLiTweets dataset and semi-automatic method for political leaning discovery
Every day, the world is flooded by millions of messages and statements posted on Twitter or Facebook. Social media platforms try to protect users' personal data, but there still is a real risk of misuse, including electi…
Detecting Multidimensional Political Incivility on Social Media
The rise of social media has been argued to intensify uncivil and hostile online political discourse. Yet, to date, there is a lack of clarity on what incivility means in the political sphere. In this work, we utilize a …
Stance Detection