paper-with-me

홈 › Papers

Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience

2024-06-03 · Martina G. Vilas, Federico Adolfi, David Poeppel, Gemma Roig

Inner Interpretability is a promising emerging field tasked with uncovering the inner mechanisms of AI systems, though how to develop these mechanistic theories is still much debated. Moreover, recent critiques raise issues that question its usefulness to advance the broader goals of AI. However, it has been overlooked that these issues resemble those that have been grappled with in another field: Cognitive Neuroscience. Here we draw the relevant connections and highlight lessons that can be transferred productively between fields. Based on these, we propose a general conceptual framework and give concrete methodological strategies for building mechanistic explanations in AI inner interpretability research. With this conceptual framework, Inner Interpretability can fend off critiques and position itself on a productive path to explain AI systems.

📄 PDF Abstract BibTeX arXiv:2406.01352

Code (0)

등록된 구현이 없습니다.

Tasks

Position

Similar Papers 제목 키워드 기반

Compositional Few-Shot Recognition with Primitive Discovery and Enhancing

2020-05-12 · Yixiong Zou, Shanghang Zhang, Ke Chen, Yonghong Tian 외

Few-shot learning (FSL) aims at recognizing novel classes given only few training samples, which still remains a great challenge for deep learning. However, humans can easily recognize novel classes with only few samples…

Few-Shot Image ClassificationFew-Shot Learningimage-classificationImage Classification+1

Analyzing Vision Transformers for Image Classification in Class Embedding Space

2023-10-29 · NeurIPS 2023 11 · Martina G. Vilas, Timothy Schaumlöffel, Gemma Roig

Despite the growing use of transformer models in computer vision, a mechanistic understanding of these networks is still needed. This work introduces a method to reverse-engineer Vision Transformers trained to solve imag…

image-classificationImage Classification

Communicability-Inspired Positional Encoding (CIPE)

2026-06-24 · Yipeng Zhang, Zhongtian Sun, Pietro Liò, Kelin Xia arxiv

Positional encodings (PEs) are essential for Transformers. Yet designing effective PEs for non-Euclidean graphs remains challenging. Such encodings should ideally induce an Attention-Compatible Geometry for self-attentio…

QIXAI: A Quantum-Inspired Framework for Enhancing Classical and Quantum Model Transparency and Understanding

2024-10-21 · John M. Willis

The impressive performance of deep learning models, particularly Convolutional Neural Networks (CNNs), is often hindered by their lack of interpretability, rendering them "black boxes." This opacity raises concerns in cr…

Feature ImportanceTime Series Analysis

Automatic Large Language Models Creation of Interactive Learning Lessons

2025-06-20 · Jionghao Lin, Jiarui Rao, Yiyang Zhao, Yuting Wang 외

We explore the automatic generation of interactive, scenario-based lessons designed to train novice human tutors who teach middle school mathematics online. Employing prompt engineering through a Retrieval-Augmented Gene…

Prompt EngineeringRetrieval-augmented Generation