paper-with-me

홈 › Papers

VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models

2026-02-03 · Woojin Kim, Sieun Hyeon, Jusang Oh, Jaeyoung Do arxiv

Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles. Value-based approaches offer a more principled path, yet three gaps persist: extraction often ignores hierarchical structure, evaluation detects presence but not calibrated intensity, and the steerability of LLMs at controlled intensities remains insufficiently understood. To address these limitations, we introduce VALUEFLOW, the first unified framework that spans extraction, evaluation, and steering with calibrated intensity control. The framework integrates three components: (i) HIVES, a hierarchical value embedding space that captures intra- and cross-theory value structure; (ii) the Value Intensity DataBase (VIDB), a large-scale resource of value-labeled texts with intensity estimates derived from ranking-based aggregation; and (iii) an anchor-based evaluator that produces consistent intensity scores for model outputs by ranking them against VIDB panels. Using VALUEFLOW, we conduct a comprehensive large-scale study across ten models and four value theories, identifying asymmetries in steerability and composition laws for multi-value control. This paper establishes a scalable infrastructure for evaluating and controlling value intensity, advancing pluralistic alignment of LLMs.

📄 PDF Abstract BibTeX arXiv:2602.03160

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression

2025-08-11 · Jadie Adams, Brian Hu, Emily Veenhuis, David Joy 외 arxiv

Large language models (LLMs) are currently aligned using techniques such as reinforcement learning from human feedback (RLHF). However, these methods use scalar rewards that can only reflect user preferences on average. …

Reinforcement Learning

Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment

2025-10-05 · Yunfan Zhang, Kathleen McKeown, Smaranda Muresan arxiv

Large Language Models (LLMs) are typically trained to reflect a relatively uniform set of values, which limits their applicability to tasks that require understanding of nuanced human perspectives. Recent research has un…

Reinforcement Learning

A Roadmap to Pluralistic Alignment

2024-02-07 · Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon 외

With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, aligning models to serve pluralistic huma…

ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems

2026-02-09 · Jinnuo Liu, Chuke Liu, Hua Shen arxiv

Multi-agent large language model (LLM) systems increasingly consist of agents that observe and respond to one another's outputs. While value alignment is typically evaluated for isolated models, how value perturbations p…

Adaptive Pluralistic Alignment: A pipeline for dynamic artificial democracy

2026-05-02 · Rachel Freedman arxiv

Prevailing alignment methods target a fixed set of preferences and therefore risk forcing value lock-in as societal norms evolve over time. We introduce Adaptive Pluralistic Alignment (APA), a modular pipeline for updati…