paper-with-me

Papers

Decipherment for Adversarial Offensive Language Detection

2018-10-01 · WS 2018 10 · Zhelun Wu, Nishant Kambhatla, Anoop Sarkar

Automated filters are commonly used by online services to stop users from sending age-inappropriate, bullying messages, or asking others to expose personal information. Previous work has focused on rules or classifiers to detect and filter offensive messages, but these are vulnerable to cleverly disguised plaintext and unseen expressions especially in an adversarial setting where the users can repeatedly try to bypass the filter. In this paper, we model the disguised messages as if they are produced by encrypting the original message using an invented cipher. We apply automatic decipherment techniques to decode the disguised malicious text, which can be then filtered using rules or classifiers. We provide experimental results on three different datasets and show that decipherment is an effective tool for this task.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DeciphermentSpelling Correction

Similar Papers 제목 키워드 기반

Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack

2019-08-17 · IJCNLP 2019 11 · Emily Dinan, Samuel Humeau, Bharath Chintagunta, Jason Weston

The detection of offensive language in the context of a dialogue has become an increasingly important application of natural language processing. The detection of trolls in public forums (Gal\'an-Garc\'ia et al., 2016), …

Sentence

Don't be a Fool: Pooling Strategies in Offensive Language Detection from User-Intended Adversarial Attacks

2024-03-20 · Seunguk Yu, Juhwan Choi, Youngbin Kim

Offensive language detection is an important task for filtering out abusive expressions and improving online user experiences. However, malicious users often attempt to avoid filtering systems through the involvement of …

Camouflage is all you need: Evaluating and Enhancing Language Model Robustness Against Camouflage Adversarial Attacks

2024-02-15 · Álvaro Huertas-García, Alejandro Martín, Javier Huertas-Tato, David Camacho

Adversarial attacks represent a substantial challenge in Natural Language Processing (NLP). This study undertakes a systematic exploration of this challenge in two distinct phases: vulnerability evaluation and resilience…

AllDecoderLanguage ModelingLanguage Modelling+1

On The Robustness of Offensive Language Classifiers

2022-03-21 · ACL 2022 5 · Jonathan Rusert, Zubair Shafiq, Padmini Srinivasan

Social media platforms are deploying machine learning based offensive language classification systems to combat hateful, racist, and other forms of offensive speech at scale. However, despite their real-world deployment,…

Decipherment-Aware Multilingual Learning in Jointly Trained Language Models

2024-06-11 · Grandee Lee

The principle that governs unsupervised multilingual learning (UCL) in jointly trained language models (mBERT as a popular example) is still being debated. Many find it surprising that one can achieve UCL with multiple m…

Decipherment