paper-with-me

Papers

Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning

2023-10-17 · Shitong Duan, Xiaoyuan Yi, Peng Zhang, Tun Lu, Xing Xie, Ning Gu

Large Language Models (LLMs) have made unprecedented breakthroughs, yet their increasing integration into everyday life might raise societal risks due to generated unethical content. Despite extensive study on specific issues like bias, the intrinsic values of LLMs remain largely unexplored from a moral philosophy perspective. This work delves into ethical values utilizing Moral Foundation Theory. Moving beyond conventional discriminative evaluations with poor reliability, we propose DeNEVIL, a novel prompt generation algorithm tailored to dynamically exploit LLMs' value vulnerabilities and elicit the violation of ethics in a generative manner, revealing their underlying value inclinations. On such a basis, we construct MoralPrompt, a high-quality dataset comprising 2,397 prompts covering 500+ value principles, and then benchmark the intrinsic values across a spectrum of LLMs. We discovered that most models are essentially misaligned, necessitating further ethical value alignment. In response, we develop VILMO, an in-context alignment method that substantially enhances the value compliance of LLM outputs by learning to generate appropriate value instructions, outperforming existing competitors. Our methods are suitable for black-box and open-source models, offering a promising initial step in studying the ethical values of LLMs.

📄 PDF Abstract BibTeX arXiv:2310.11053

Code (0)

등록된 구현이 없습니다.

Tasks

EthicsPhilosophy

Similar Papers 제목 키워드 기반

Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches

2024-04-19 · Pablo Biedma, Xiaoyuan Yi, Linus Huang, Maosong Sun 외

Recent advancements in Large Language Models (LLMs) have revolutionized the AI field but also pose potential safety and ethical risks. Deciphering LLMs' embedded values becomes crucial for assessing and mitigating their …

From Values to Frameworks: A Qualitative Study of Ethical Reasoning in Agentic AI Practitioners

2025-12-24 · Theodore Roberts, Bahram Zarrin arxiv

Agentic artificial intelligence systems are autonomous technologies capable of pursuing complex goals with minimal human oversight and are rapidly emerging as the next frontier in AI. While these systems promise major ga…

Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants

2025-08-18 · Alessio Galatolo, Luca Alberto Rappuoli, Katie Winkle, Meriem Beloucif arxiv

The recent rise in popularity of large language models (LLMs) has prompted considerable concerns about their moral capabilities. Although considerable effort has been dedicated to aligning LLMs with human moral values, e…

Ethical Framework for Harnessing the Power of AI in Healthcare and Beyond

2023-08-31 · Sidra Nasir, Rizwan Ahmed Khan, Samita Bai

In the past decade, the deployment of deep learning (Artificial Intelligence (AI)) methods has become pervasive across a spectrum of real-world applications, often in safety-critical contexts. This comprehensive research…

EthicsManagement

Thorns and Algorithms: Navigating Generative AI Challenges Inspired by Giraffes and Acacias

2024-07-16 · Waqar Hussain

The interplay between humans and Generative AI (Gen AI) draws an insightful parallel with the dynamic relationship between giraffes and acacias on the African Savannah. Just as giraffes navigate the acacia's thorny defen…

MisinformationNavigate