paper-with-me

홈 › Papers

DFKI SLT at GermEval 2021: Multilingual Pre-training and Data Augmentation for the Classification of Toxicity in Social Media Comments

2021-09-01 · GermEval 2021 9 · Remi Calizzano, Malte Ostendorff, Georg Rehm

We present our submission to the first subtask of GermEval 2021 (classification of German Facebook comments as toxic or not). Binary sequence classification is a standard NLP task with known state-of-the-art methods. Therefore, we focus on data preparation by using two different techniques: task-specific pre-training and data augmentation. First, we pre-train multilingual transformers (XLM-RoBERTa and MT5) on 12 hatespeech detection datasets in nine different languages. In terms of F1, we notice an improvement of 10% on average, using task-specific pre-training. Second, we perform data augmentation by labelling unlabelled comments, taken from Facebook, to increase the size of the training dataset by 79%. Models trained on the augmented training dataset obtain on average +0.0282 (+5%) F1 score compared to models trained on the original training dataset. Finally, the combination of the two techniques allows us to obtain an F1 score of 0.6899 with XLM- RoBERTa and 0.6859 with MT5. The code of the project is available at: https://github.com/airKlizz/germeval2021toxic.

📄 PDF Abstract BibTeX

Code (1)

airklizz/germeval2021toxic 공식 구현 pytorch

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

AIT_FHSTP at GermEval 2021: Automatic Fact Claiming Detection with Multilingual Transformer Models

2021-09-01 · GermEval 2021 9 · Jaqueline Böck, Daria Liakhovets, Mina Schütz, Armin Kirchknopf 외

Spreading ones opinion on the internet is becoming more and more important. A problem is that in many discussions people often argue with supposed facts. This year’s GermEval 2021 focuses on this topic by incorporating a…

AIxcellent Vibes at GermEval 2025 Shared Task on Candy Speech Detection: Improving Model Performance by Span-Level Training

2025-09-09 · Christian Rene Thelen, Patrick Gustav Blaneck, Tobias Bornheim, Niklas Grieger 외 arxiv

Positive, supportive online communication in social media (candy speech) has the potential to foster civility, yet automated detection of such language remains underexplored, limiting systematic analysis of its impact. W…

ur-iw-hnt at GermEval 2021: An Ensembling Strategy with Multiple BERT Models

2021-10-05 · GermEval 2021 9 · Hoai Nam Tran, Udo Kruschwitz

This paper describes our approach (ur-iw-hnt) for the Shared Task of GermEval2021 to identify toxic, engaging, and fact-claiming comments. We submitted three runs using an ensembling strategy by majority (hard) voting wi…

Going beyond zero-shot MT: combining phonological, morphological and semantic factors. The UdS-DFKI System at IWSLT 2017

2017-12-01 · IWSLT 2017 12 · Cristina España-Bonet, Josef van Genabith

This paper describes the UdS-DFKI participation to the multilingual task of the IWSLT Evaluation 2017. Our approach is based on factored multilingual neural translation systems following the small data and zero-shot trai…

Translation

Precog-LTRC-IIITH at GermEval 2021: Ensembling Pre-Trained Language Models with Feature Engineering

2021-09-01 · GermEval 2021 9 · T. H. Arjun, Arvindh A., Kumaraguru Ponnurangam

We describe our participation in all the subtasks of the Germeval 2021 shared task on the identification of Toxic, Engaging, and Fact-Claiming Comments. Our system is an ensemble of state-of-the-art pre-trained models fi…

Data AugmentationFeature Engineering