paper-with-me

Papers

Concept Alignment as a Prerequisite for Value Alignment

2023-10-30 · Sunayana Rane, Mark Ho, Ilia Sucholutsky, Thomas L. Griffiths

Value alignment is essential for building AI systems that can safely and reliably interact with people. However, what a person values -- and is even capable of valuing -- depends on the concepts that they are currently using to understand and evaluate what happens in the world. The dependence of values on concepts means that concept alignment is a prerequisite for value alignment -- agents need to align their representation of a situation with that of humans in order to successfully align their values. Here, we formally analyze the concept alignment problem in the inverse reinforcement learning setting, show how neglecting concept alignment can lead to systematic value mis-alignment, and describe an approach that helps minimize such failure modes by jointly reasoning about a person's concepts and values. Additionally, we report experimental results with human participants showing that humans reason about the concepts used by an agent when acting intentionally, in line with our joint reasoning model.

📄 PDF Abstract BibTeX arXiv:2310.20059

Code (0)

등록된 구현이 없습니다.

Tasks

Concept Alignment

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers

2024-12-09 · Johanna Vielhaben, Dilyara Bareeva, Jim Berend, Wojciech Samek 외

Vision transformers (ViTs) can be trained using various learning paradigms, from fully supervised to self-supervised. Diverse training protocols often result in significantly different feature spaces, which are usually c…

CLLMRec: LLM-powered Cognitive-Aware Concept Recommendation via Semantic Alignment and Prerequisite Knowledge Distillation

2025-11-21 · Xiangrui Xiong, Yichuan Lu, Zifei Pan, Chang Sun arxiv

The growth of Massive Open Online Courses (MOOCs) presents significant challenges for personalized learning, where concept recommendation is crucial. Existing approaches typically rely on heterogeneous information networ…

Knowledge DistillationKnowledge TracingKnowledge Graphs

Concept Alignment

2024-01-09 · Sunayana Rane, Polyphony J. Bruna, Ilia Sucholutsky, Christopher Kello 외

Discussion of AI alignment (alignment between humans and AI systems) has focused on value alignment, broadly referring to creating AI systems that share human values. We argue that before we can even attempt to align val…

Concept AlignmentPhilosophy

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

2026-09-04 · Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans arxiv

AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there is no correct target, but a shared prerequisite is that the system's behavior expresses a coherent policy: a m…

ConTrans: Weak-to-Strong Alignment Engineering via Concept Transplantation

2024-05-22 · Weilong Dong, Xinwei Wu, Renren Jin, Shaoyang Xu 외

Ensuring large language models (LLM) behave consistently with human goals, values, and intentions is crucial for their safety but yet computationally expensive. To reduce the computational cost of alignment training of L…