paper-with-me

Papers

Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models

2022-06-16 · Maribeth Rauh, John Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, Iason Gabriel, William Isaac, Lisa Anne Hendricks

Large language models produce human-like text that drive a growing number of applications. However, recent literature and, increasingly, real world observations, have demonstrated that these models can generate language that is toxic, biased, untruthful or otherwise harmful. Though work to evaluate language model harms is under way, translating foresight about which harms may arise into rigorous benchmarks is not straightforward. To facilitate this translation, we outline six ways of characterizing harmful text which merit explicit consideration when designing new benchmarks. We then use these characteristics as a lens to identify trends and gaps in existing benchmarks. Finally, we apply them in a case study of the Perspective API, a toxicity classifier that is widely used in harm benchmarks. Our characteristics provide one piece of the bridge that translates between foresight and effective evaluation.

📄 PDF Abstract BibTeX arXiv:2206.08325

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingLanguage ModelingLanguage ModellingTranslation

Similar Papers 제목 키워드 기반

Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models

2025-09-05 · Youjia Zheng, Mohammad Zandsalimy, Shanu Sushmita arxiv

Large Language Models (LLMs) are increasingly vulnerable to a sophisticated form of adversarial prompting known as camouflaged jailbreaking. This method embeds malicious intent within seemingly benign language to evade e…

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

2025-02-03 · Zora Che, Stephen Casper, Robert Kirk, Anirudh Satheesh 외

Evaluations of large language model (LLM) risks and capabilities are increasingly being incorporated into AI risk management and governance frameworks. Currently, most risk evaluations are conducted by designing inputs t…

BenchmarkingLarge Language Model

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

2026-06-25 · Jiajun Wu, Haoyu Kang, Yining Sun, Jiacheng Hou 외 arxiv

Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks. However, we identify two primary limitations i…

Binary Classification

Immunization against harmful fine-tuning attacks

2024-02-26 · Domenic Rosati, Jan Wehner, Kai Williams, Łukasz Bartoszcze 외

Large Language Models (LLMs) are often trained with safety guards intended to prevent harmful text generation. However, such safety training can be removed by fine-tuning the LLM on harmful datasets. While this emerging …

Text Generation

IMGTB: A Framework for Machine-Generated Text Detection Benchmarking

2023-11-21 · Michal Spiegel, Dominik Macko

In the era of large language models generating high quality texts, it is a necessity to develop methods for detection of machine-generated text to avoid harmful use or simply due to annotation purposes. It is, however, a…

BenchmarkingText Detection