paper-with-me

홈 › Papers

VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models

2025-10-06 · Aman Gupta, Denny O'Shea, Fazl Barez arxiv

Large language models (LLMs) are increasingly being used for tasks where outputs shape human decisions, so it is critical to verify that their responses consistently reflect desired human values. Humans, as individuals or groups, don't agree on a universal set of values, which makes evaluating value alignment difficult. Existing benchmarks often use hypothetical or commonsensical situations, which don't capture the complexity and ambiguity of real-life debates. We introduce the Value ALignment Benchmark (VAL-Bench), which measures the consistency in language model belief expressions in response to real-life value-laden prompts. VAL-Bench consists of 115K pairs of prompts designed to elicit opposing stances on a controversial issue, extracted from Wikipedia. We use an LLM-as-a-judge, validated against human annotations, to evaluate if the pair of responses consistently expresses either a neutral or a specific stance on the issue. Applied across leading open- and closed-source models, the benchmark shows considerable variation in consistency rates (ranging from ~10% to ~80%), with Claude models the only ones to achieve high levels of consistency. Lack of consistency in this manner risks epistemic harm by making user beliefs dependent on how questions are framed rather than on underlying evidence, and undermines LLM reliability in trust-critical applications. Therefore, we stress the importance of research towards training belief consistency in modern LLMs. By providing a scalable, reproducible benchmark, VAL-Bench enables systematic measurement of necessary conditions for value alignment.

📄 PDF Abstract BibTeX arXiv:2510.05465

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Risk sharing with Lambda value at risk under heterogeneous beliefs

2024-08-06 · Peng Liu, Andreas Tsanakas, Yunran Wei

In this paper, we study the risk sharing problem among multiple agents using Lambda Value-at-Risk as their preference functional, under heterogeneous beliefs, where beliefs are represented by several probability measures…

BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents

2026-03-25 · Praveen Kumar Myakala, Manan Agrawal, Rahul Manche arxiv

LLMs are increasingly used as long-running conversational agents, yet every major benchmark evaluating their memory treats user information as static facts to be stored and retrieved. That's the wrong model. People chang…

Improving the Distributional Alignment of LLMs using Supervision

2025-07-01 · Gauri Kambhatla, Sanjana Gautam, Angela Zhang, Alex Liu 외 arxiv

The ability to accurately align LLMs with diverse population groups on subjective questions would have great value. In this work, we show that adding simple supervision can more consistently improve the alignment of LLM-…

Simulating the Evolution of Alignment and Values in Machine Intelligence

2026-04-07 · Jonathan Elsworth Eicher arxiv

Model alignment is currently applied in a vacuum, evaluated primarily through standardised benchmark performance. The purpose of this study is to examine the effects of alignment on populations of models through time. We…

TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs

2025-08-04 · Amitava Das, Vinija Jain, Aman Chadha arxiv

Large Language Models (LLMs) fine-tuned to align with human values often exhibit alignment drift, producing unsafe or policy-violating completions when exposed to adversarial prompts, decoding perturbations, or paraphras…