paper-with-me

Papers

EigenBench: A Comparative Behavioral Measure of Value Alignment

2025-09-02 · Jonathn Chang, Leonhard Piff, Suvadip Sana, Jasmine X. Li, Lionel Levine arxiv

Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for comparatively benchmarking language models' values. Given an ensemble of models, a constitution describing a value system, and a dataset of scenarios, our method returns a vector of scores quantifying each model's alignment to the given constitution. To produce these scores, each model judges the outputs of other models across many scenarios, and these judgments are aggregated with EigenTrust (Kamvar et al., 2003), yielding scores that reflect a weighted consensus judgment of the whole ensemble. EigenBench uses no ground truth labels, as it is designed to quantify subjective traits for which reasonable judges may disagree on the correct label. Hence, to validate our method, we collect human judgments on the same ensemble of models and show that EigenBench's judgments align closely with those of human evaluators. We further demonstrate that EigenBench can recover model rankings on the GPQA benchmark without access to objective labels, supporting its viability as a framework for evaluating subjective values for which no ground truths exist. The code is available at https://github.com/jchang153/EigenBench.

📄 PDF Abstract BibTeX arXiv:2509.01938

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating Representational Similarity Measures from the Lens of Functional Correspondence

2024-11-21 · Yiqing Bo, Ansh Soni, Sudhanshu Srivastava, Meenakshi Khosla

Neuroscience and artificial intelligence (AI) both face the challenge of interpreting high-dimensional neural data, where the comparative analysis of such data is crucial for revealing shared mechanisms and differences b…

Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models

2025-11-19 · Samih Fadli arxiv

Large language model safety is usually assessed with static benchmarks, but key failures are dynamic: value drift under distribution shift, jailbreak attacks, and slow degradation of alignment in deployment. Building on …

Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without a Ground-Truth Standard

2026-06-07 · Alireza Arbabi, Florian Kerschbaum arxiv

Large language models (LLMs) are increasingly released and deployed through opaque development and deployment pipelines, enabling model providers to inject intentional, provider-specific policies without officially annou…

Adaptive Hardness-driven Augmentation and Alignment Strategies for Multi-Source Domain Adaptations

2025-01-02 · Yang Yuxiang, Zeng Xinyi, Zeng Pinxian, Zu Chen 외

Multi-source Domain Adaptation (MDA) aims to transfer knowledge from multiple labeled source domains to an unlabeled target domain. Nevertheless, traditional methods primarily focus on achieving inter-domain alignment th…

Data AugmentationDomain Adaptation

Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment

2025-08-27 · Julian Arnold, Niels Lörch arxiv

Fine-tuning LLMs on narrowly harmful datasets can lead to behavior that is broadly misaligned with respect to human values. To understand when and how this emergent misalignment occurs, we develop a comprehensive framewo…

Change Detection