paper-with-me

홈 › Papers

ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations

2024-06-18 · Yunze Xiao, Yujia Hu, Kenny Tsu Wei Choo, Roy Ka-Wei Lee

Detecting hate speech and offensive language is essential for maintaining a safe and respectful digital environment. This study examines the limitations of state-of-the-art large language models (LLMs) in identifying offensive content within systematically perturbed data, with a focus on Chinese, a language particularly susceptible to such perturbations. We introduce \textsf{ToxiCloakCN}, an enhanced dataset derived from ToxiCN, augmented with homophonic substitutions and emoji transformations, to test the robustness of LLMs against these cloaking perturbations. Our findings reveal that existing models significantly underperform in detecting offensive content when these perturbations are applied. We provide an in-depth analysis of how different types of offensive content are affected by these perturbations and explore the alignment between human and model explanations of offensiveness. Our work highlights the urgent need for more advanced techniques in offensive language detection to combat the evolving tactics used to evade detection mechanisms.

📄 PDF Abstract BibTeX arXiv:2406.12223

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection

2024-10-21 · Jianfei He, Lilin Wang, Jiaying Wang, Zhenyu Liu 외

Identifying offensive language is essential for maintaining safety and sustainability in the social media era. Though large language models (LLMs) have demonstrated encouraging potential in social media analytics, they l…

Camouflage is all you need: Evaluating and Enhancing Language Model Robustness Against Camouflage Adversarial Attacks

2024-02-15 · Álvaro Huertas-García, Alejandro Martín, Javier Huertas-Tato, David Camacho

Adversarial attacks represent a substantial challenge in Natural Language Processing (NLP). This study undertakes a systematic exploration of this challenge in two distinct phases: vulnerability evaluation and resilience…

AllDecoderLanguage ModelingLanguage Modelling+1

Detection of Offensive and Threatening Online Content in a Low Resource Language

2023-11-17 · Fatima Muhammad Adam, Abubakar Yakubu Zandam, Isa Inuwa-Dutse

Hausa is a major Chadic language, spoken by over 100 million people in Africa. However, from a computational linguistic perspective, it is considered a low-resource language, with limited resources to support Natural Lan…

AttributeTranslation

Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack

2019-08-17 · IJCNLP 2019 11 · Emily Dinan, Samuel Humeau, Bharath Chintagunta, Jason Weston

The detection of offensive language in the context of a dialogue has become an increasingly important application of natural language processing. The detection of trolls in public forums (Gal\'an-Garc\'ia et al., 2016), …

Sentence

Hate-Speech and Offensive Language Detection in Roman Urdu

2020-11-01 · EMNLP 2020 11 · Hammad Rizwan, Muhammad Haroon Shakeel, Asim Karim

The task of automatic hate-speech and offensive language detection in social media content is of utmost importance due to its implications in unprejudiced society concerning race, gender, or religion. Existing research i…

Transfer Learning