paper-with-me

홈 › Papers

Navigating the Concept Space of Language Models

2026-03-06 · Wilson E. Marcílio-Jr, Danilo M. Eler arxiv

Sparse autoencoders (SAEs) trained on large language model activations output thousands of features that enable mapping to human-interpretable concepts. The current practice for analyzing these features primarily relies on inspecting top-activating examples, manually browsing individual features, or performing semantic search on interested concepts, which makes exploratory discovery of concepts difficult at scale. In this paper, we present Concept Explorer, a scalable interactive system for post-hoc exploration of SAE features that organizes concept explanations using hierarchical neighborhood embeddings. Our approach constructs a multi-resolution manifold over SAE feature embeddings and enables progressive navigation from coarse concept clusters to fine-grained neighborhoods, supporting discovery, comparison, and relationship analysis among concepts. We demonstrate the utility of Concept Explorer on SAE features extracted from SmolLM2, where it reveals coherent high-level structure, meaningful subclusters, and distinctive rare concepts that are hard to identify with existing workflows.

📄 PDF Abstract BibTeX arXiv:2603.23524

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Navigating the Conceptual Multiverse

2026-04-20 · Andre Ye, Jenny Y. Huang, Alicia Guo, Rose Novick 외 arxiv

When language models answer open-ended problems, they implicitly make hidden decisions that shape their outputs, leaving users with uncontextualized answers rather than a working map of the problem; drawing on multiverse…

Navigating Conceptual Space; A new take on Artificial General Intelligence

2022-02-19 · Per R. Leikanger

Edward C. Tolman found reinforcement learning unsatisfactory for explaining intelligence and proposed a clear distinction between learning and behavior. Tolman's ideas on latent learning and cognitive maps eventually led…

Autonomous NavigationRobot Navigationvalid

Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence

2022-02-07 · Frederik Pahde, Maximilian Dreyer, Leander Weber, Moritz Weckbecker 외

With a growing interest in understanding neural network prediction strategies, Concept Activation Vectors (CAVs) have emerged as a popular tool for modeling human-understandable concepts in the latent space. Commonly, CA…

TAG

AI-as-exploration: Navigating intelligence space

2024-01-15 · Dimitri Coelho Mollo

Artificial Intelligence is a field that lives many lives, and the term has come to encompass a motley collection of scientific and commercial endeavours. In this paper, I articulate the contours of a rather neglected but…

Mixture of Concept Bottleneck Experts

2026-02-02 · Francesco De Santis, Gabriele Ciravegna, Giovanni De Felice, Arianna Casanova 외 arxiv

Concept Bottleneck Models (CBMs) promote interpretability by grounding predictions in human-understandable concepts. However, existing CBMs typically constrain their task predictor to a single expression whose functional…