paper-with-me

홈 › Papers

Mirror-Neuron Patterns in AI Alignment

2025-10-23 · Robyn Wyrick arxiv

As artificial intelligence (AI) advances toward superhuman capabilities, aligning these systems with human values becomes increasingly critical. Current alignment strategies rely largely on externally specified constraints that may prove insufficient against future super-intelligent AI capable of circumventing top-down controls. This research investigates whether artificial neural networks (ANNs) can develop patterns analogous to biological mirror neurons cells that activate both when performing and observing actions, and how such patterns might contribute to intrinsic alignment in AI. Mirror neurons play a crucial role in empathy, imitation, and social cognition in humans. The study therefore asks: (1) Can simple ANNs develop mirror-neuron patterns? and (2) How might these patterns contribute to ethical and cooperative decision-making in AI systems? Using a novel Frog and Toad game framework designed to promote cooperative behaviors, we identify conditions under which mirror-neuron patterns emerge, evaluate their influence on action circuits, introduce the Checkpoint Mirror Neuron Index (CMNI) to quantify activation strength and consistency, and propose a theoretical framework for further study. Our findings indicate that appropriately scaled model capacities and self/other coupling foster shared neural representations in ANNs similar to biological mirror neurons. These empathy-like circuits support cooperative behavior and suggest that intrinsic motivations modeled through mirror-neuron dynamics could complement existing alignment techniques by embedding empathy-like mechanisms directly within AI architectures.

📄 PDF Abstract BibTeX arXiv:2511.01885

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Embodied Representation Alignment with Mirror Neurons

2025-09-25 · Wentao Zhu, Zhining Zhang, Yuwei Ren, Yin Huang 외 arxiv

Mirror neurons are a class of neurons that activate both when an individual observes an action and when they perform the same action. This mechanism reveals a fundamental interplay between action understanding and embodi…

Representation LearningAction UnderstandingContrastive Learning

The Mysterious Case of Neuron 1512: Injectable Realignment Architectures Reveal Internal Characteristics of Meta's Llama 2 Model

2024-07-04 · Brenden Smith, Dallin Baker, Clayton Chase, Myles Barney 외

Large Language Models (LLMs) have an unrivaled and invaluable ability to "align" their output to a diverse range of human preferences, by mirroring them in the text they generate. The internal characteristics of such mod…

Language ModelingLanguage Modelling

Brain-like Functional Organization within Large Language Models

2024-10-25 · Haiyang Sun, Lin Zhao, Zihao Wu, Xiaohui Gao 외

The human brain has long inspired the pursuit of artificial intelligence (AI). Recently, neuroimaging studies provide compelling evidence of alignment between the computational representation of artificial neural network…

Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding

2025-10-19 · Yudan Ren, Xinlong Wang, Kexin Wang, Tian Xia 외 arxiv

While brain-inspired artificial intelligence(AI) has demonstrated promising results, current understanding of the parallels between artificial neural networks (ANNs) and human brain processing remains limited: (1) unimod…

Mirror Neuron; A Beautiful Unnecessary Concept

2019-11-14

The mirror neuron theory that has enjoyed continued validations was developed with no particular attention to the phenomenon of the vision. Understandably the perception of vision has always been thought to happen, natur…