paper-with-me

홈 › Papers

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

2026-08-12 · Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen hf

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.

📄 PDF Abstract BibTeX arXiv:2608.12036

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Open Problems in Mechanistic Interpretability

2025-01-27 · Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey 외

Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goals. Progress in this field thus promises…

Position: Prioritize Identifying Structure, Not Complex Models, for Scientific Discovery

2026-05-30 · Tyler H. McCormick arxiv

Modern Machine Learning (ML) and Artificial Intelligence (AI) models, especially large language models (LLMs), are increasingly used to generate scientific hypotheses and mechanistic explanations from observational data.…

Instrumented data for causal scientific machine learning

2026-06-05 · Daniel N. Wilke arxiv

Scientific machine learning is limited less by model size than by the data it is trained on. Observational data records what happened but not why; template synthetic data has a known generating process but only for the s…

A Probabilistic Framework for LLM-Based Model Discovery

2026-02-20 · Stefan Wahl, Raphaela Schenk, Ali Farnoud, Jakob H. Macke 외 arxiv

Automated methods for discovering mechanistic simulator models from observational data offer a promising path toward accelerating scientific progress. Such methods often take the form of agentic-style iterative workflows…

The Quest for an Integrated Set of Neural Mechanisms Underlying Object Recognition in Primates

2023-12-10 · Kohitij Kar, James J DiCarlo

Visual object recognition -- the behavioral ability to rapidly and accurately categorize many visually encountered objects -- is core to primate cognition. This behavioral capability is algorithmically impressive because…

ObjectObject Recognition