paper-with-me

홈 › Papers

AIxcellent Vibes at GermEval 2025 Shared Task on Candy Speech Detection: Improving Model Performance by Span-Level Training

2025-09-09 · Christian Rene Thelen, Patrick Gustav Blaneck, Tobias Bornheim, Niklas Grieger, Stephan Bialonski arxiv

Positive, supportive online communication in social media (candy speech) has the potential to foster civility, yet automated detection of such language remains underexplored, limiting systematic analysis of its impact. We investigate how candy speech can be reliably detected in a 46k-comment German YouTube corpus by monolingual and multilingual language models, including GBERT, Qwen3 Embedding, and XLM-RoBERTa. We find that a multilingual XLM-RoBERTa-Large model trained to detect candy speech at the span level outperforms other approaches, ranking first in both binary positive F1: 0.8906) and categorized span-based detection (strict F1: 0.6307) subtasks at the GermEval 2025 Shared Task on Candy Speech Detection. We speculate that span-based training, multilingual capabilities, and emoji-aware tokenizers improved detection performance. Our results demonstrate the effectiveness of multilingual models in identifying positive, supportive language.

📄 PDF Abstract BibTeX arXiv:2509.07459

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Overview of the GermEval 2021 Shared Task on the Identification of Toxic, Engaging, and Fact-Claiming Comments

2021-09-01 · GermEval 2021 9 · Julian Risch, Anke Stoll, Lena Wilms, Michael Wiegand

We present the GermEval 2021 shared task on the identification of toxic, engaging, and fact-claiming comments. This shared task comprises three binary classification subtasks with the goal to identify: toxic comments, en…

Binary ClassificationFact Checking

WLV-RIT at GermEval 2021: Multitask Learning with Transformers to Detect Toxic, Engaging, and Fact-Claiming Comments

2021-07-30 · GermEval 2021 9 · Skye Morgan, Tharindu Ranasinghe, Marcos Zampieri

This paper addresses the identification of toxic, engaging, and fact-claiming comments on social media. We used the dataset made available by the organizers of the GermEval-2021 shared task containing over 3,000 manually…

AIT_FHSTP at GermEval 2021: Automatic Fact Claiming Detection with Multilingual Transformer Models

2021-09-01 · GermEval 2021 9 · Jaqueline Böck, Daria Liakhovets, Mina Schütz, Armin Kirchknopf 외

Spreading ones opinion on the internet is becoming more and more important. A problem is that in many discussions people often argue with supposed facts. This year’s GermEval 2021 focuses on this topic by incorporating a…

IRCologne at GermEval 2021: Toxicity Classification

2021-09-01 · GermEval 2021 9 · Fabian Haak, Björn Engelmann

In this paper, we describe the TH Köln’s submission for the Shared Task on the Identification of Toxic Comments at GermEval 2021. Toxicity is a severe and latent problem in comments in online discussions. Complex languag…

ClassificationLanguage ModelingLanguage ModellingToxic Comment Classification

TUW-Inf at GermEval2021: Rule-based and Hybrid Methods for Detecting Toxic, Engaging, and Fact-Claiming Comments

2021-09-01 · GermEval 2021 9 · Kinga Gémes, Gábor Recski

This paper describes our methods submitted for the GermEval 2021 shared task on identifying toxic, engaging and fact-claiming comments in social media texts (Risch et al., 2021). We explore simple strategies for semi-aut…