paper-with-me

홈 › Papers

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

2026-01-20 · Hengyuan Zhang, Zhihao Zhang, Mingyang Wang, Zunhai Su, Yiwei Wang, Qianli Wang, Shuzhou Yuan, Ercong Nie, Xufeng Duan, Feijiang Han, Qibo Xue, Zeping Yu, Chenming Shang, Xiao Liang, Jing Xiong, Hui Shen, Chaofan Tao, Zhengwu Liu, Senjie Jin, Zhiheng Xi, Dongdong Zhang, Sophia Ananiadou, Tao Gui, Ruobing Xie, Hayden Kwok-Hay So, Hinrich Schütze, Xuanjing Huang, Qi Zhang, Ngai Wong arxiv

Mechanistic Interpretability (MI) has emerged as a vital approach to demystify the opaque decision-making of Large Language Models (LLMs). However, existing reviews primarily treat MI as an observational science, summarizing analytical insights while lacking a systematic framework for actionable intervention. To bridge this gap, we present a practical survey structured around the pipeline: "Locate, Steer, and Improve." We formally categorize Localizing (diagnosis) and Steering (intervention) methods based on specific Interpretable Objects to establish a rigorous intervention protocol. Furthermore, we demonstrate how this framework enables tangible improvements in Alignment, Capability, and Efficiency, effectively operationalizing MI as an actionable methodology for model optimization. The curated paper list of this work is available at https://github.com/rattlesnakey/Awesome-Actionable-MI-Survey.

📄 PDF Abstract BibTeX arXiv:2601.14004

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Attribution to Action: A Human-Centered Application of Activation Steering

2026-04-13 · Tobias Labarta, Maximilian Dreyer, Katharina Weitz, Wojciech Samek 외 arxiv

Explainable AI (XAI) methods reveal which features influence model predictions, yet provide limited means for practitioners to act on these explanations. Activation steering of components identified via XAI offers a path…

Erasing Concepts, Steering Generations: A Comprehensive Survey of Concept Suppression

2025-05-26 · Yiwei Xie, Ping Liu, Zheng Zhang

Text-to-Image (T2I) models have demonstrated impressive capabilities in generating high-quality and diverse visual content from natural language prompts. However, uncontrolled reproduction of sensitive, copyrighted, or h…

Adversarial RobustnessDisentanglementSpecificity

SQAPlanner: Generating Data-Informed Software Quality Improvement Plans

2021-02-19 · Dilini Rajapaksha, Chakkrit Tantithamthavorn, Jirayus Jiarpakdee, Christoph Bergmeir 외

Software Quality Assurance (SQA) planning aims to define proactive plans, such as defining maximum file size, to prevent the occurrence of software defects in future releases. To aid this, defect prediction models have b…

SAEs Are Good for Steering -- If You Select the Right Features

2025-05-26 · Dana Arad, Aaron Mueller, Yonatan Belinkov

Sparse Autoencoders (SAEs) have been proposed as an unsupervised approach to learn a decomposition of a model's latent space. This enables useful applications such as steering - influencing the output of a model towards …

Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey

2022-10-14 · Sachin Kumar, Vidhisha Balachandran, Lucille Njoo, Antonios Anastasopoulos 외

Recent advances in the capacity of large language models to generate human-like text have resulted in their increased adoption in user-facing settings. In parallel, these improvements have prompted a heated discourse aro…

Language ModelingLanguage ModellingSurveyText Generation