paper-with-me

홈 › Papers

Radical AI Interpretability

2026-06-25 · Daniel A. Herrmann, Benjamin A. Levinstein arxiv

We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The core question is: given the computational facts about a system, how do we solve for its beliefs, desires, and meanings? This matters increasingly for safety. We want to be able to trust the systems we deploy, whether by understanding their goals or, more modestly, by reliably detecting deception. Interpretability researchers are building tools to read beliefs and desires off a model's internals, but there is no settled account of when such a tool has succeeded. This book supplies one. We propose criteria on both representationalist and interpretationist approaches, and tie each to tests current interpretability methods can carry out. A central lesson is that these attributions cannot be made piecemeal. Beliefs, desires, and the propositional structure they presuppose are jointly constrained, and a method that fixes one while measuring the others inherits whatever distortions that introduces. This holism becomes pressing for AI systems, which may not share the interpreter's concepts. However, it also provides leverage: a system's attitudes constrain its propositional structure, that structure constrains which attitudes can be attributed, and mechanistic interpretability can help us measure both.

📄 PDF Abstract BibTeX arXiv:2606.26523

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Radically Compositional Cognitive Concepts

2019-11-14 · Toby B. St Clere Smithe

Despite ample evidence that our concepts, our cognitive architecture, and mathematics itself are all deeply compositional, few models take advantage of this structure. We therefore propose a radically compositional appro…

Interpretable Oracle Bone Script Decipherment through Radical and Pictographic Analysis with LVLMs

2025-08-13 · Kaixin Peng, Mengyang Zhao, Haiyang Yu, Teng Fu 외 arxiv

As the oldest mature writing system, Oracle Bone Script (OBS) has long posed significant challenges for archaeological decipherment due to its rarity, abstractness, and pictographic diversity. Current deep learning-based…

Radical-Enhanced Chinese Character Embedding

2014-04-18 · Yaming Sun, Lei Lin, Duyu Tang, Nan Yang 외

We present a method to leverage radical for learning Chinese character embedding. Radical is a semantic and phonetic component of Chinese character. It plays an important role as characters with the same radical usually …

Chinese Word Segmentation

Chinese Character Recognition with Radical-Structured Stroke Trees

2022-11-24 · Haiyang Yu, Jingye Chen, Bin Li, xiangyang xue

The flourishing blossom of deep learning has witnessed the rapid development of Chinese character recognition. However, it remains a great challenge that the characters for testing may have different distributions from t…

Decoder

Understanding the Radical Mind: Identifying Signals to Detect Extremist Content on Twitter

2019-05-15 · Mariam Nouh, Jason R. C. Nurse, Michael Goldsmith

The Internet and, in particular, Online Social Networks have changed the way that terrorist and extremist groups can influence and radicalise individuals. Recent reports show that the mode of operation of these groups st…