paper-with-me

홈 › Papers

Explaining Toxic Text via Knowledge Enhanced Text Generation

2022-07-01 · NAACL 2022 7 · Rohit Sridhar, Diyi Yang

Warning: This paper contains content that is offensive and may be upsetting.Biased or toxic speech can be harmful to various demographic groups. Therefore, it is not only important for models to detect these speech, but to also output explanations of why a given text is toxic. Previous literature has mostly focused on classifying and detecting toxic speech, and existing efforts on explaining stereotypes in toxic speech mainly use standard text generation approaches, resulting in generic and repetitive explanations. Building on these prior works, we introduce a novel knowledge-informed encoder-decoder framework to utilize multiple knowledge sources to generate implications of biased text.Experiments show that our knowledge informed models outperform prior state-of-the-art models significantly, and can generate detailed explanations of stereotypes in toxic speech compared to baselines, both quantitatively and qualitatively.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderText Generation

Similar Papers 제목 키워드 기반

Explaining Toxic Text via Knowledge Enhanced Text Generation

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Biased or toxic speech can be harmful to various demographic groups. Therefore, it is not only important for models to detect these speech, but to also output explanations of why a given text is toxic. Previous literatur…

DecoderText Generation

Just KIDDIN: Knowledge Infusion and Distillation for Detection of INdecent Memes

2024-11-19 · Rahul Garg, Trilok Padhi, Hemang Jain, Ugur Kursuncu 외

Toxicity identification in online multimodal environments remains a challenging task due to the complexity of contextual connections across modalities (e.g., textual and visual). In this paper, we propose a novel framewo…

Knowledge DistillationKnowledge Graphs

Toxicity Detection can be Sensitive to the Conversational Context

2021-11-19 · Alexandros Xenos, John Pavlopoulos, Ion Androutsopoulos, Lucas Dixon 외

User posts whose perceived toxicity depends on the conversational context are rare in current toxicity detection datasets. Hence, toxicity detectors trained on existing datasets will also tend to disregard context, makin…

Data AugmentationKnowledge Distillation

Simulated Ensemble Attack: Transferring Jailbreaks Across Fine-tuned Vision-Language Models

2025-08-03 · Ruofan Wang, Xin Wang, Yang Yao, Juncheng Li 외 arxiv

The widespread practice of fine-tuning open-source Vision-Language Models (VLMs) raises a critical security concern: jailbreak vulnerabilities in base models may persist in downstream variants, enabling transferable atta…

Text Detoxification as Style Transfer in English and Hindi

2024-02-12 · Sourabrata Mukherjee, Akanksha Bansal, Atul Kr. Ojha, John P. McCrae 외

This paper focuses on text detoxification, i.e., automatically converting toxic text into non-toxic text. This task contributes to safer and more respectful online communication and can be considered a Text Style Transfe…

Multi-Task LearningSentenceStyle TransferText Style Transfer+1