paper-with-me

홈 › Papers

No Offense Taken: Eliciting Offensiveness from Language Models

2023-10-02 · Anugya Srivastava, Rahul Ahuja, Rohith Mukku

This work was completed in May 2022. For safe and reliable deployment of language models in the real world, testing needs to be robust. This robustness can be characterized by the difficulty and diversity of the test cases we evaluate these models on. Limitations in human-in-the-loop test case generation has prompted an advent of automated test case generation approaches. In particular, we focus on Red Teaming Language Models with Language Models by Perez et al.(2022). Our contributions include developing a pipeline for automated test case generation via red teaming that leverages publicly available smaller language models (LMs), experimenting with different target LMs and red classifiers, and generating a corpus of test cases that can help in eliciting offensive responses from widely deployed LMs and identifying their failure modes.

📄 PDF Abstract BibTeX arXiv:2310.00892

Code (1)

anugyas/nluproject 공식 구현

Tasks

DiversityRed Teaming

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

FII FUNNY at SemEval-2021 Task 7: HaHackathon: Detecting and rating Humor and Offense

2021-08-01 · SEMEVAL 2021 · Mihai Samson, Daniela Gifu

The {``}HaHackathon: Detecting and Rating Humor and Offense{''} task at the SemEval 2021 competition focuses on detecting and rating the humor level in sentences, as well as the level of offensiveness contained in these …

Humor@IITK at SemEval-2021 Task 7: Large Language Models for Quantifying Humor and Offensiveness

2021-04-02 · SEMEVAL 2021 · Aishwarya Gupta, Avik Pal, Bholeshwar Khurana, Lakshay Tyagi 외

Humor and Offense are highly subjective due to multiple word senses, cultural knowledge, and pragmatic competence. Hence, accurately detecting humorous and offensive texts has several compelling use cases in Recommendati…

Recommendation Systems

Smatgrisene at SemEval-2020 Task 12: Offense Detection by AI - with a Pinch of Real I

2020-12-01 · SEMEVAL 2020 · Peter Juel Henrichsen, Marianne Rathje

This paper discusses how ML based classifiers can be enhanced disproportionately by adding small amounts of qualitative linguistic knowledge. As an example we present the Danish classifier Smatgrisene, our contribution t…

IIITG-ADBU at SemEval-2020 Task 12: Comparison of BERT and BiLSTM in Detecting Offensive Language

2020-12-01 · SEMEVAL 2020 · Arup Baruah, Kaushik Das, Ferdous Barbhuiya, Kuntal Dey

Task 12 of SemEval 2020 consisted of 3 subtasks, namely offensive language identification (Subtask A), categorization of offense type (Subtask B), and offense target identification (Subtask C). This paper presents the re…

Language IdentificationWorld Knowledge

DLJUST at SemEval-2021 Task 7: Hahackathon: Linking Humor and Offense

2021-08-01 · SEMEVAL 2021 · Hani Al-Omari, Isra{'}a AbedulNabi, Rehab Duwairi

Humor detection and rating poses interesting linguistic challenges to NLP; it is highly subjective depending on the perceptions of a joke and the context in which it is used. This paper utilizes and compares transformers…

Humor Detection