paper-with-me

Papers

Concept Alignment

2024-01-09 · Sunayana Rane, Polyphony J. Bruna, Ilia Sucholutsky, Christopher Kello, Thomas L. Griffiths

Discussion of AI alignment (alignment between humans and AI systems) has focused on value alignment, broadly referring to creating AI systems that share human values. We argue that before we can even attempt to align values, it is imperative that AI systems and humans align the concepts they use to understand the world. We integrate ideas from philosophy, cognitive science, and deep learning to explain the need for concept alignment, not just value alignment, between humans and machines. We summarize existing accounts of how humans and machines currently learn concepts, and we outline opportunities and challenges in the path towards shared concepts. Finally, we explain how we can leverage the tools already being developed in cognitive science and AI research to accelerate progress towards concept alignment.

📄 PDF Abstract BibTeX arXiv:2401.08672

Code (0)

등록된 구현이 없습니다.

Tasks

Concept AlignmentPhilosophy

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Concept Alignment as a Prerequisite for Value Alignment

2023-10-30 · Sunayana Rane, Mark Ho, Ilia Sucholutsky, Thomas L. Griffiths

Value alignment is essential for building AI systems that can safely and reliably interact with people. However, what a person values -- and is even capable of valuing -- depends on the concepts that they are currently u…

Concept Alignment

Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers

2024-12-09 · Johanna Vielhaben, Dilyara Bareeva, Jim Berend, Wojciech Samek 외

Vision transformers (ViTs) can be trained using various learning paradigms, from fully supervised to self-supervised. Diverse training protocols often result in significantly different feature spaces, which are usually c…

Probing the Probes: Methods and Metrics for Concept Alignment

2025-11-06 · Jacob Lysnæs-Larsen, Marte Eggen, Inga Strümke arxiv

In explainable AI, Concept Activation Vectors (CAVs) are typically obtained by training linear classifier probes to detect human-understandable concepts as directions in the activation space of deep neural networks. It i…

A Taxonomy of Conceptual Alignment in Human-Robot Dialogue

2026-06-21 · Shengchen Zhang, Xiaohua Sun, Weiwei Guo arxiv

Successful conversations require speakers to align on the meaning of concepts, a challenging but crucial task for human-robot interaction. Understanding the process of establishing such alignment is hindered by competing…

Concept-wise Attention for Fine-grained Concept Bottleneck Models

2026-04-17 · Minghong Zhong, Guoshuai Zou, Kanghao Chen, Dexia Chen 외 arxiv

Recently impressive performance has been achieved in Concept Bottleneck Models (CBM) by utilizing the image-text alignment learned by a large pre-trained vision-language model (i.e. CLIP). However, there exist two key li…