paper-with-me

Papers

From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach

2026-05-20 · Nura Aljaafari, Danilo S. Carvalho, Andre Freitas arxiv

Mechanistic interpretability produces circuit-level causal analyses of neural network behaviour, but discovered circuits often remain isolated experimental artefacts: there is no shared formal representation for what circuits compute, how they relate, or when two findings provide evidence for the same mechanism. This work provides a formal infrastructure for cumulative mechanistic science by treating circuit interpretation as inductive theory construction. Each circuit is characterised at two levels: a Causal Functional Signature (CFS), which grounds component behaviour in causal attribution evidence and token role profiles, and an architectural signature $τ_{\mathrm{arch}}$, learned by inductive logic programming (ILP) from scale-invariant structural predicates. Together, these constitute a formal coherence layer that makes mechanistic claims explicit, comparable via $θ$-subsumption, and portable across model scales. CFS reveals qualitatively distinct computational strategies across task types, including attention-mediated copying versus MLP-mediated binding. ILP signatures achieve substantially better structural separation than graph kernel and feature-vector baselines, and support principled transfer across model scales and architecture families.

📄 PDF Abstract BibTeX arXiv:2605.21303

Code (0)

등록된 구현이 없습니다.

Tasks

Inductive logic programming

Similar Papers 제목 키워드 기반

Neural mechanisms underlying the temporal organization of naturalistic animal behavior

2022-03-04 · Luca Mazzucato

Naturalistic animal behavior exhibits a strikingly complex organization in the temporal domain, whose variability stems from at least three sources: hierarchical, contextual, and stochastic. What are the neural mechanism…

Towards a Mechanistic Understanding of Propositional Logical Reasoning in Large Language Models

2026-01-07 · Danchun Chen, Qiyao Yan, Liangming Pan arxiv

Understanding how Large Language Models (LLMs) perform logical reasoning internally remains a fundamental challenge. While prior mechanistic studies focus on identifying taskspecific circuits, they leave open the questio…

Logical Reasoning

A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations

2023-02-06 · Bilal Chughtai, Lawrence Chan, Neel Nanda

Universality is a key hypothesis in mechanistic interpretability -- that different models learn similar features and circuits when trained on similar tasks. In this work, we study the universality hypothesis by examining…

Mechanistic Foundations of Goal-Directed Control

2026-03-16 · Alma Lago arxiv

Mechanistic interpretability has transformed the analysis of transformer circuits by decomposing model behavior into competing algorithms, identifying phase transitions during training, and deriving closed-form predictio…

How Transformers Solve Propositional Logic Problems: A Mechanistic Analysis

2024-11-06 · Guan Zhe Hong, Nishanth Dikkala, Enming Luo, Cyrus Rashtchian 외

Large language models (LLMs) have shown amazing performance on tasks that require planning and reasoning. Motivated by this, we investigate the internal mechanisms that underpin a network's ability to perform complex log…

Logical Reasoning