paper-with-me

홈 › Papers

SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models

2025-11-25 · Jiaojiao Han, Wujiang Xu, Mingyu Jin, Mengnan Du arxiv

Large language models (LLMs) have achieved remarkable progress, yet their internal mechanisms remain largely opaque, posing a significant challenge to their safe and reliable deployment. Sparse autoencoders (SAEs) have emerged as a promising tool for decomposing LLM representations into more interpretable features, but explaining the features captured by SAEs remains a challenging task. In this work, we propose SAGE (SAE AGentic Explainer), an agent-based framework that recasts feature interpretation from a passive, single-pass generation task into an active, explanation-driven process. SAGE implements a rigorous methodology by systematically formulating multiple explanations for each feature, designing targeted experiments to test them, and iteratively refining explanations based on empirical activation feedback. Experiments on features from SAEs of diverse language models demonstrate that SAGE produces explanations with significantly higher generative and predictive accuracy compared to state-of-the-art baselines.an agent-based framework that recasts feature interpretation from a passive, single-pass generation task into an active, explanationdriven process. SAGE implements a rigorous methodology by systematically formulating multiple explanations for each feature, designing targeted experiments to test them, and iteratively refining explanations based on empirical activation feedback. Experiments on features from SAEs of diverse language models demonstrate that SAGE produces explanations with significantly higher generative and predictive accuracy compared to state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2511.20820

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interpreting GNN-based IDS Detections Using Provenance Graph Structural Features

2023-06-01 · Kunal Mukherjee, Joshua Wiedemeier, Tianhao Wang, Muhyun Kim 외

Advanced cyber threats (e.g., Fileless Malware and Advanced Persistent Threat (APT)) have driven the adoption of provenance-based security solutions. These solutions employ Machine Learning (ML) models for behavioral mod…

Anomaly DetectionDecision MakingDescriptiveExplainable Models+3

SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization

2026-06-07 · Jingyi He, Haiyan Zhao, Ruxue Shi, Yanguang Liu 외 arxiv

Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse features, explaining these features still remains a central challenge. Curre…

Parameterized Explainer for Graph Neural Network

2020-11-09 · NeurIPS 2020 12 · Dongsheng Luo, Wei Cheng, Dongkuan Xu, Wenchao Yu 외

Despite recent progress in Graph Neural Networks (GNNs), explaining predictions made by GNNs remains a challenging open problem. The leading method independently addresses the local explanations (i.e., important subgraph…

Graph ClassificationGraph Neural Network

CoRTX: Contrastive Framework for Real-time Explanation

2023-03-05 · Yu-Neng Chuang, Guanchu Wang, Fan Yang, Quan Zhou 외

Recent advancements in explainable machine learning provide effective and faithful solutions for interpreting model behaviors. However, many explanation methods encounter efficiency issues, which largely limit their depl…

Contrastive Learning

LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models

2024-06-21 · Mengdan Zhu, Raasikh Kanjiani, Jiahui Lu, Andrew Choi 외

Deep generative models like VAEs and diffusion models have advanced various generation tasks by leveraging latent variables to learn data distributions and generate high-quality samples. Despite the field of explainable …

Uncertainty Quantification