An Affective-Taxis Hypothesis for Alignment and Interpretability
AI alignment is a field of research that aims to develop methods to ensure that agents always behave in a manner aligned with (i.e. consistently with) the goals and values of their human operators, no matter their level of capability. This paper proposes an affectivist approach to the alignment problem, re-framing the concepts of goals and values in terms of affective taxis, and explaining the emergence of affective valence by appealing to recent work in evolutionary-developmental and computational neuroscience. We review the state of the art and, building on this work, we propose a computational model of affect based on taxis navigation. We discuss evidence in a tractable model organism that our model reflects aspects of biological taxis navigation. We conclude with a discussion of the role of affective taxis in AI alignment.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Latent Structure of Affective Representations in Large Language Models
The geometric structure of latent representations in large language models (LLMs) is an active area of research, driven in part by its implications for model transparency and AI safety. Existing literature has focused ma…
Towards Interpretability in Audio and Visual Affective Machine Learning: A Review
Machine learning is frequently used in affective computing, but presents challenges due the opacity of state-of-the-art machine learning methods. Because of the impact affective machine learning systems may have on an in…
ArticlesDecision MakingExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
Large language models (LLMs) have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque. Mechanistic interpretability (i.e., the systematic study of how…
Reinforcement LearningPlay with Emotion: Affect-Driven Reinforcement Learning
This paper introduces a paradigm shift by viewing the task of affect modeling as a reinforcement learning (RL) process. According to the proposed paradigm, RL agents learn a policy (i.e. affective interaction) by attempt…
Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
Large language models appear to develop internal representations of emotion -- "emotion circuits," "emotion neurons," and structured emotional manifolds have been reported across multiple model families. But every study …