paper-with-me

홈 › Papers

An Affective-Taxis Hypothesis for Alignment and Interpretability

2025-05-03 · Eli Sennesh, Maxwell Ramstead

AI alignment is a field of research that aims to develop methods to ensure that agents always behave in a manner aligned with (i.e. consistently with) the goals and values of their human operators, no matter their level of capability. This paper proposes an affectivist approach to the alignment problem, re-framing the concepts of goals and values in terms of affective taxis, and explaining the emergence of affective valence by appealing to recent work in evolutionary-developmental and computational neuroscience. We review the state of the art and, building on this work, we propose a computational model of affect based on taxis navigation. We discuss evidence in a tractable model organism that our model reflects aspects of biological taxis navigation. We conclude with a discussion of the role of affective taxis in AI alignment.

📄 PDF Abstract BibTeX arXiv:2505.17024

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Latent Structure of Affective Representations in Large Language Models

2026-04-08 · Benjamin J. Choi, Melanie Weber arxiv

The geometric structure of latent representations in large language models (LLMs) is an active area of research, driven in part by its implications for model transparency and AI safety. Existing literature has focused ma…

Towards Interpretability in Audio and Visual Affective Machine Learning: A Review

2023-06-15 · David S. Johnson, Olya Hakobyan, Hanna Drimalla

Machine learning is frequently used in affective computing, but presents challenges due the opacity of state-of-the-art machine learning methods. Because of the impact affective machine learning systems may have on an in…

ArticlesDecision MakingExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)

Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions

2026-01-21 · Usman Naseem arxiv

Large language models (LLMs) have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque. Mechanistic interpretability (i.e., the systematic study of how…

Reinforcement Learning

Play with Emotion: Affect-Driven Reinforcement Learning

2022-08-26 · Matthew Barthet, Ahmed Khalifa, Antonios Liapis, Georgios N. Yannakakis

This paper introduces a paradigm shift by viewing the task of affect modeling as a reinforcement learning (RL) process. According to the proposed paradigm, RL agents learn a policy (i.e. affective interaction) by attempt…

Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs

2026-03-15 · Michael Keeman arxiv

Large language models appear to develop internal representations of emotion -- "emotion circuits," "emotion neurons," and structured emotional manifolds have been reported across multiple model families. But every study …