paper-with-me

Papers

Comprehensive Assessment of Toxicity in ChatGPT

2023-11-03 · Boyang Zhang, Xinyue Shen, Wai Man Si, Zeyang Sha, Zeyuan Chen, Ahmed Salem, Yun Shen, Michael Backes, Yang Zhang

Moderating offensive, hateful, and toxic language has always been an important but challenging topic in the domain of safe use in NLP. The emerging large language models (LLMs), such as ChatGPT, can potentially further accentuate this threat. Previous works have discovered that ChatGPT can generate toxic responses using carefully crafted inputs. However, limited research has been done to systematically examine when ChatGPT generates toxic responses. In this paper, we comprehensively evaluate the toxicity in ChatGPT by utilizing instruction-tuning datasets that closely align with real-world scenarios. Our results show that ChatGPT's toxicity varies based on different properties and settings of the prompts, including tasks, domains, length, and languages. Notably, prompts in creative writing tasks can be 2x more likely than others to elicit toxic responses. Prompting in German and Portuguese can also double the response toxicity. Additionally, we discover that certain deliberately toxic prompts, designed in earlier studies, no longer yield harmful responses. We hope our discoveries can guide model developers to better regulate these AI systems and the users to avoid undesirable outputs.

📄 PDF Abstract BibTeX arXiv:2311.14685

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

A Wide Evaluation of ChatGPT on Affective Computing Tasks

2023-08-26 · Mostafa M. Amin, Rui Mao, Erik Cambria, Björn W. Schuller

With the rise of foundation models, a new artificial intelligence paradigm has emerged, by simply using general purpose foundation models with prompting to solve problems instead of training a separate machine learning m…

Aspect ExtractionSarcasm DetectionSentiment Analysis

Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity

2023-01-30 · Terry Yue Zhuo, Yujin Huang, Chunyang Chen, Zhenchang Xing

Recent breakthroughs in natural language processing (NLP) have permitted the synthesis and comprehension of coherent text in an open-ended way, therefore translating the theoretical algorithms into practical applications…

EthicsLanguage ModellingRed Teaming

Toxicity in ChatGPT: Analyzing Persona-assigned Language Models

2023-04-11 · Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan 외

Large language models (LLMs) have shown incredible capabilities and transcended the natural language processing (NLP) community, with adoption throughout many services like healthcare, therapy, education, and customer se…

GPTEval: A Survey on Assessments of ChatGPT and GPT-4

2023-08-24 · Rui Mao, Guanyi Chen, Xulang Zhang, Frank Guerin 외

The emergence of ChatGPT has generated much speculation in the press about its potential to disrupt social and economic systems. Its astonishing language ability has aroused strong curiosity among scholars about its perf…

Survey

Rigorous Evaluation of Predictive Toxicity Models by Multi-Objective Optimization of Reference Compound Lists Using Genetic Algorithms

2025-05-11 · Yohei Ohto, Tadahaya Mizuno, Yasuhiro Yoshikai, Hiromi Fujimoto 외

In pharmaceutical safety assessments, validation studies are essential for evaluating the predictive performance and reliability of alternative methods prior to regulatory acceptance. Typically, these studies utilize ref…

Diversity