paper-with-me

홈 › Papers

Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms

2026-06-06 · Hyunjin Cho, Youngji Roh, Jaehyung Kim arxiv

As large language models are increasingly deployed in high-stakes settings, there is a growing need for tools that audit not only model outputs but also the internal computations that produce them. Circuit analysis is a central approach in mechanistic interpretability, but it is typically target-conditioned, explaining a single prompt paired with a chosen completion. This target-conditioned setup can obscure heterogeneity across a model's continuation distribution. We introduce distribution-level unsupervised feature discovery, which clusters sampled continuations using both semantic content and sequence-level mechanistic attributions, without manually specifying target outputs. Our method represents each continuation with a semantic embedding and a prefix-to-continuation attribution signature, then optimizes a rate-distortion objective that trades off semantic coherence, mechanistic consistency, and cluster granularity. Across clustering and steering analyses, the discovered clusters expose continuation modes that single-view baselines miss and provide interventional evidence that cluster signatures correspond to actionable mechanistic factors. Overall, our approach complements circuit analysis and behavioral evaluation by providing a scalable audit of the mechanisms underlying a model's continuation distribution.

📄 PDF Abstract BibTeX arXiv:2606.08236

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Unsupervised Cross-domain Image-to-Image Translation Using a Shared Discriminator

2021-02-09 · Rajiv Kumar, Rishabh Dabral, G. Sivakumar

Unsupervised image-to-image translation is used to transform images from a source domain to generate images in a target domain without using source-target image pairs. Promising results have been obtained for this proble…

Image-to-Image TranslationTranslationUnsupervised Image-To-Image Translation

Unsupervised Feature Learning through Divergent Discriminative Feature Accumulation

2014-06-06 · Paul A. Szerlip, Gregory Morse, Justin K. Pugh, Kenneth O. Stanley

Unlike unsupervised approaches such as autoencoders that learn to reconstruct their inputs, this paper introduces an alternative approach to unsupervised feature learning called divergent discriminative feature accumulat…

General Classification

Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation

2018-09-06 · CONLL 2018 10 · Mikel Artetxe, Gorka Labaka, Iñigo Lopez-Gazpio, Eneko Agirre

Following the recent success of word embeddings, it has been argued that there is no such thing as an ideal representation for words, as different models tend to capture divergent and often mutually incompatible aspects …

Word Embeddings

Divergent Ensemble Networks: Enhancing Uncertainty Estimation with Shared Representations and Independent Branching

2024-12-02 · Arnav Kharbanda, Advait Chandorkar

Ensemble learning has proven effective in improving predictive performance and estimating uncertainty in neural networks. However, conventional ensemble methods often suffer from redundant parameter usage and computation…

DiversityEnsemble LearningRepresentation Learning

LDRE: LLM-based Divergent Reasoning and Ensemble for Zero-Shot Composed Image Retrieval

2024-07-11 · SIGIR 2024 7 · Zhenyu Yang, Dizhan Xue, Shengsheng Qian, WeiMing Dong 외

Zero-Shot Composed Image Retrieval (ZS-CIR) has garnered increasing interest in recent years, which aims to retrieve a target image based on a query composed of a reference image and a modification text without training …

Image RetrievalImage to textRetrievalZero-Shot Composed Image Retrieval (ZS-CIR)