White-Box Attacks on Hate-speech BERT Classifiers in German with Explicit and Implicit Character Level Defense
In this work, we evaluate the adversarial robustness of BERT models trained on German Hate Speech datasets. We also complement our evaluation with two novel white-box character and word level attacks thereby contributing to the range of attacks available. Furthermore, we also perform a comparison of two novel character-level defense strategies and evaluate their robustness with one another.
Code (1)
Tasks
Adversarial RobustnessMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Hate Speech Detection and Racial Bias Mitigation in Social Media based on BERT model
Disparate biases associated with datasets and trained classifiers in hateful and abusive content identification tasks have raised many concerns recently. Although the problem of biased datasets on abusive language detect…
Abusive LanguageHate Speech DetectionLanguage ModellingTransfer LearningDetecting White Supremacist Hate Speech using Domain Specific Word Embedding with Deep Learning and BERT
White supremacists embrace a radical ideology that considers white people superior to people of other races. The critical influence of these groups is no longer limited to social media; they also have a significant effec…
Language ModellingTowards non-toxic landscapes: Automatic toxic comment detection using DNN
The spectacular expansion of the Internet has led to the development of a new research problem in the field of natural language processing: automatic toxic comment detection, since many countries prohibit hate speech in …
Binary ClassificationCharacter-level HyperNetworks for Hate Speech Detection
The massive spread of hate speech, hateful content targeted at specific subpopulations, is a problem of critical social importance. Automated methods of hate speech detection typically employ state-of-the-art deep learni…
Data AugmentationHate Speech DetectionAre Chess Discussions Racist? An Adversarial Hate Speech Data Set
On June 28, 2020, while presenting a chess podcast on Grandmaster Hikaru Nakamura, Antonio Radi\'c's YouTube handle got blocked because it contained "harmful and dangerous" content. YouTube did not give further specific …