paper-with-me

홈 › Papers

CounterMoral: Editing Morals in Language Models

2026-03-28 · Michael Ripa, Jim Davies arxiv

Recent advancements in language model technology have significantly enhanced the ability to edit factual information. Yet, the modification of moral judgments, a crucial aspect of aligning models with human values, has garnered less attention. In this work, we introduce CounterMoral, a benchmark dataset crafted to assess how well current model editing techniques modify moral judgments across diverse ethical frameworks. We apply various editing techniques to multiple language models and evaluate their performance. Our findings contribute to the evaluation of language models designed to be ethical.

📄 PDF Abstract BibTeX arXiv:2603.27338

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

V-MORALS: Visual Morse Graph-Aided Estimation of Regions of Attraction in a Learned Latent Space

2026-02-26 · Faiz Aladin, Ashwin Balasubramanian, Lars Lindemann, Daniel Seita arxiv

Reachability analysis has become increasingly important in robotics to distinguish safe from unsafe states. Unfortunately, existing reachability and safety analysis methods often fall short, as they typically require kno…

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

2025-05-25 · Steffen Backmann, David Guzman Piedrahita, Emanuel Tewolde, Rada Mihalcea 외

Recent advances in large language models (LLMs) have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a key AI safety concern. While prior work h…

EthicsNavigate

A Corpus for Understanding and Generating Moral Stories

2022-04-20 · NAACL 2022 7 · Jian Guan, Ziqi Liu, Minlie Huang

Teaching morals is one of the most important purposes of storytelling. An essential ability for understanding and writing moral stories is bridging story plots and implied morals. Its challenges mainly lie in: (1) graspi…

Retrieval

MoVa: Towards Generalizable Classification of Human Morals and Values

2025-09-29 · Ziyu Chen, Junfei Sun, Chenxi Li, Tuan Dung Nguyen 외 arxiv

Identifying human morals and values embedded in language is essential to empirical studies of communication. However, researchers often face substantial difficulty navigating the diversity of theoretical frameworks and d…

Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation

2026-04-09 · Sophie Wu, Andrew Piper arxiv

Stories are key to transmitting values across cultures, but their interpretation varies across linguistic and cultural contexts. Thus, we introduce multilingual story moral generation as a novel culturally grounded evalu…

Semantic Similarity