paper-with-me

홈 › Papers

PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm

2026-01-13 · Jing-Jing Li, Joel Mire, Eve Fleisig, Valentina Pyatkin, Anne Collins, Maarten Sap, Sydney Levine arxiv

Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond consensus and instead understand where and why disagreements arise. We introduce PluriHarms, a benchmark designed to systematically study human harm judgments across two key dimensions -- the harm axis (benign to harmful) and the agreement axis (agreement to disagreement). Our scalable framework generates prompts that capture diverse AI harms and human values while targeting cases with high disagreement rates, validated by human data. The benchmark includes 150 prompts with 15,000 ratings from 100 human annotators, enriched with demographic and psychological traits and prompt-level features of harmful actions, effects, and values. Our analyses show that prompts that relate to imminent risks and tangible harms amplify perceived harmfulness, while annotator traits (e.g., toxicity experience, education) and their interactions with prompt content explain systematic disagreement. We benchmark AI safety models and alignment methods on PluriHarms, finding that while personalization significantly improves prediction of human harm judgments, considerable room remains for future progress. By explicitly targeting value diversity and disagreement, our work provides a principled benchmark for moving beyond "one-size-fits-all" safety toward pluralistically safe AI.

📄 PDF Abstract BibTeX arXiv:2601.08951

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Perspectives on Large Language Models for Relevance Judgment

2023-04-13 · Guglielmo Faggioli, Laura Dietz, Charles Clarke, Gianluca Demartini 외

When asked, large language models (LLMs) like ChatGPT claim that they can assist with relevance judgments but it is not clear whether automated judgments can reliably be used in evaluations of retrieval systems. In this …

Retrieval

Towards a Gold Standard for Evaluating Danish Word Embeddings

2020-05-01 · LREC 2020 5 · Nina Schneidermann, Rasmus Hvingelby, Bolette Pedersen

This paper presents the process of compiling a model-agnostic similarity goal standard for evaluating Danish word embeddings based on human judgments made by 42 native speakers of Danish. Word embeddings resemble semanti…

Semantic SimilaritySemantic Textual SimilarityWord Embeddings

STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions

2024-09-20 · Robert Morabito, Sangmitra Madhusudan, Tyler McDonald, Ali Emami

Mitigating explicit and implicit biases in Large Language Models (LLMs) has become a critical focus in the field of natural language processing. However, many current methodologies evaluate scenarios in isolation, withou…

BenchmarkingSensitivity

Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments

2026-01-12 · Bingyang Ye, Shan Chen, Jingxuan Tu, Chen Liu 외 arxiv

Large language models are increasingly being used to assess and forecast research ideas, yet we lack scalable ways to evaluate the quality of models' judgments about these scientific ideas. Towards this goal, we introduc…

Advancing Annotation of Stance in Social Media Posts: A Comparative Analysis of Large Language Models and Crowd Sourcing

2024-06-11 · Mao Li, Frederick Conrad

In the rapidly evolving landscape of Natural Language Processing (NLP), the use of Large Language Models (LLMs) for automated text annotation in social media posts has garnered significant interest. Despite the impressiv…

BenchmarkingStance Detectiontext annotation