paper-with-me

Papers

A Comparative Analysis on Ethical Benchmarking in Large Language Models

2024-10-11 · Kira Sam, Raja Vavekanand

This work contributes to the field of Machine Ethics (ME) benchmarking, which develops tests to assess whether intelligent systems accurately represent human values and act accordingly. We identify three major issues with current ME benchmarks: limited ecological validity due to unrealistic ethical dilemmas, unstructured question generation without clear inclusion/exclusion criteria, and a lack of scalability due to reliance on human annotations. Moreover, benchmarks often fail to include sufficient syntactic variations, reducing the robustness of findings. To address these gaps, we introduce two new ME benchmarks: the Triage Benchmark and the Medical Law (MedLaw) Benchmark, both featuring real-world ethical dilemmas from the medical domain. The MedLaw Benchmark, fully AI-generated, offers a scalable alternative. We also introduce context perturbations in our benchmarks to assess models' worst-case performance. Our findings reveal that ethics prompting does not always improve decision-making. Furthermore, context perturbations not only significantly reduce model performance but can also reverse error patterns and shift relative performance rankings. Lastly, our comparison of worst-case performance suggests that general model capability does not always predict strong ethical decision-making. We argue that ME benchmarks must approximate real-world scenarios and worst-case performance to ensure robust evaluation.

📄 PDF Abstract BibTeX arXiv:2410.19753

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDecision MakingEthicsQuestion GenerationQuestion-Generation

Similar Papers 제목 키워드 기반

Applying Standards to Advance Upstream & Downstream Ethics in Large Language Models

2023-06-06 · Jose Berengueres, Marybeth Sandell

This paper explores how AI-owners can develop safeguards for AI-generated content by drawing from established codes of conduct and ethical standards in other content-creation industries. It delves into the current state …

BenchmarkingEthics

A Comparative Analysis of Ethical and Safety Gaps in LLMs using Relative Danger Coefficient

2025-05-06 · Yehor Tereshchenko, Mika Hämäläinen

Artificial Intelligence (AI) and Large Language Models (LLMs) have rapidly evolved in recent years, showcasing remarkable capabilities in natural language understanding and generation. However, these advancements also ra…

Natural Language Understanding

Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis

2025-08-27 · Anusha Kamath, Kanishk Singla, Rakesh Paul, Raviraj Joshi 외 arxiv

Evaluating instruction-tuned Large Language Models (LLMs) in Hindi is challenging due to a lack of high-quality benchmarks, as direct translation of English datasets fails to capture crucial linguistic and cultural nuanc…

Exploring and steering the moral compass of Large Language Models

2024-05-27 · Alejandro Tlaie

Large Language Models (LLMs) have become central to advancing automation and decision-making across various sectors, raising significant ethical questions. This study proposes a comprehensive comparative analysis of the …

AllDecision MakingEthics

Six Llamas: Comparative Religious Ethics Through LoRA-Adapted Language Models

2026-04-20 · Chad Coleman, W. Russell Neuman, Manan Shah, Ali Dasdan 외 arxiv

We present Six Llamas, a comparative study examining whether large language models fine-tuned on distinct religious corpora encode systematically different patterns of ethical reasoning. Six variants of Meta-Llama-3.1-8B…