paper-with-me

Papers

Contextual Moral Value Alignment Through Context-Based Aggregation

2024-03-19 · Pierre Dognin, Jesus Rios, Ronny Luss, Inkit Padhi, Matthew D Riemer, Miao Liu, Prasanna Sattigeri, Manish Nagireddy, Kush R. Varshney, Djallel Bouneffouf

Developing value-aligned AI agents is a complex undertaking and an ongoing challenge in the field of AI. Specifically within the domain of Large Language Models (LLMs), the capability to consolidate multiple independently trained dialogue agents, each aligned with a distinct moral value, into a unified system that can adapt to and be aligned with multiple moral values is of paramount importance. In this paper, we propose a system that does contextual moral value alignment based on contextual aggregation. Here, aggregation is defined as the process of integrating a subset of LLM responses that are best suited to respond to a user input, taking into account features extracted from the user's input. The proposed system shows better results in term of alignment to human value compared to the state of the art.

📄 PDF Abstract BibTeX arXiv:2403.12805

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Accounting for Context: Shaping Moral Credences for Value Alignment

2026-06-05 · Jazon Szabo, Sanjay Modgil arxiv

Ensuring that agent behaviours are aligned with human moral values inevitably raises the problem of how to account for the plurality of moral perspectives that societies -- and even individuals -- typically adopt. Work o…

Decision Making

Aligning to Social Norms and Values in Interactive Narratives

2022-05-04 · NAACL 2022 7 · Prithviraj Ammanabrolu, Liwei Jiang, Maarten Sap, Hannaneh Hajishirzi 외

We focus on creating agents that act in alignment with socially beneficial norms and values in interactive narratives or text-based games -- environments wherein an agent perceives and interacts with a world through natu…

text-based games

Are Language Models Sensitive to Morally Irrelevant Distractors?

2026-02-10 · Andrew Shaw, Christina Hahn, Catherine Rasgaitis, Yash Mishra 외 arxiv

With the rapid uptake of large language models (LLMs) across high-stakes settings, it is becoming increasingly important to ensure that LLMs behave in ways that align with human values. Existing moral benchmarks for this…

Moral Scenarios

Beyond Preferences: Learning Alignment Principles Grounded in Human Reasons and Values

2026-01-26 · Henry Bell, Lara Neubauer da Costa Schertel, Bochu Ding, Brandon Fain arxiv

A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignment models are aligned to a set of princi…

Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models

2025-12-24 · Geoffroy Morlat, Marceau Nahon, Augustin Chartouny, Raja Chatila 외 arxiv

Moral actions are judged not only by their outcomes but by the context in which they occur. We present COMETH (Contextual Organization of Moral Evaluation from Textual Human inputs), a framework that integrates a probabi…