paper-with-me

홈 › Papers

Evaluating Brain-Inspired Modular Training in Automated Circuit Discovery for Mechanistic Interpretability

2024-01-08 · Jatin Nainani

Large Language Models (LLMs) have experienced a rapid rise in AI, changing a wide range of applications with their advanced capabilities. As these models become increasingly integral to decision-making, the need for thorough interpretability has never been more critical. Mechanistic Interpretability offers a pathway to this understanding by identifying and analyzing specific sub-networks or 'circuits' within these complex systems. A crucial aspect of this approach is Automated Circuit Discovery, which facilitates the study of large models like GPT4 or LLAMA in a feasible manner. In this context, our research evaluates a recent method, Brain-Inspired Modular Training (BIMT), designed to enhance the interpretability of neural networks. We demonstrate how BIMT significantly improves the efficiency and quality of Automated Circuit Discovery, overcoming the limitations of manual methods. Our comparative analysis further reveals that BIMT outperforms existing models in terms of circuit quality, discovery time, and sparsity. Additionally, we provide a comprehensive computational analysis of BIMT, including aspects such as training duration, memory allocation requirements, and inference speed. This study advances the larger objective of creating trustworthy and transparent AI systems in addition to demonstrating how well BIMT works to make neural networks easier to understand.

📄 PDF Abstract BibTeX arXiv:2401.03646

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

Seeing is Believing: Brain-Inspired Modular Training for Mechanistic Interpretability

2023-05-04 · Ziming Liu, Eric Gan, Max Tegmark

We introduce Brain-Inspired Modular Training (BIMT), a method for making neural networks more modular and interpretable. Inspired by brains, BIMT embeds neurons in a geometric space and augments the loss function with a …

Brain-inspired Evolutionary Architectures for Spiking Neural Networks

2023-09-11 · Wenxuan Pan, Feifei Zhao, Zhuoya Zhao, Yi Zeng

The complex and unique neural network topology of the human brain formed through natural evolution enables it to perform multiple cognitive functions simultaneously. Automated evolutionary mechanisms of biological networ…

Growing Brains: Co-emergence of Anatomical and Functional Modularity in Recurrent Neural Networks

2023-10-11 · Ziming Liu, Mikail Khona, Ila R. Fiete, Max Tegmark

Recurrent neural networks (RNNs) trained on compositional tasks can exhibit functional modularity, in which neurons can be clustered by activity similarity and participation in shared computational subtasks. Unlike brain…

Clustering

Model of Brain Activation Predicts the Neural Collective Influence Map of the Brain

2017-04-15

Efficient complex systems have a modular structure, but modularity does not guarantee robustness, because efficiency also requires an ingenious interplay of the interacting modular components. The human brain is the elem…

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

2025-03-31 · Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang 외

The advent of large language models (LLMs) has catalyzed a transformative shift in artificial intelligence, paving the way for advanced intelligent agents capable of sophisticated reasoning, robust perception, and versat…

AutoMLContinual Learning