paper-with-me

Papers

The Information Geometry of Softmax: Probing and Steering

2026-02-17 · Kiho Park, Todd Nief, Yo Joong Choe, Victor Veitch arxiv

This paper concerns the question of how AI systems encode semantic structure into the geometric structure of their representation spaces. The motivating observation is that the natural geometry of these representation spaces should reflect the way models use representations to produce behavior. We focus on the important special case of representations that define softmax distributions. In this case, we argue that the natural geometry is information geometry. Our focus is on the role of information geometry on semantic encoding and the linear representation hypothesis. As an illustrative application, we develop "dual steering", a method for robustly steering representations to exhibit a particular concept using linear probes. We prove that dual steering optimally modifies the target concept while minimizing changes to off-target concepts. Empirically, we find that dual steering enhances the controllability and stability of concept manipulation.

📄 PDF Abstract BibTeX arXiv:2602.15293

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stream separation improves Bregman conditioning in transformers

2026-03-22 · James Clayton Kerce arxiv

Linear methods for steering transformer representations, including probing, activation engineering, and concept erasure, implicitly assume the geometry of representation space is Euclidean. Park et al. [Park et al., 2026…

The Geometry of Harmfulness in LLMs through Subconcept Probing

2025-07-23 · McNair Shah, Saleena Angeline, Adhitya Rajendra Kumar, Naitik Chheda 외 arxiv

Recent advances in large language models (LLMs) have intensified the need to understand and reliably curb their harmful behaviours. We introduce a multidimensional framework for probing and steering harmful content in mo…

FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers

2026-05-17 · Sihan Wang, Jiayi Zhao arxiv

Activation steering has emerged as a lightweight approach for modifying language model behavior without parameter updates, yet existing methods remain brittle: unstable across layers and prone to disturbing behavior unre…

A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

2026-07-01 · Siyi Wang, James Bailey, Ting Dang arxiv

While prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly understood. We present the first comparati…

Speech Synthesis

Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States

2026-06-01 · Subramanyam Sahoo, Vinija Jain, Aman Chadha, Divya Chaudhary arxiv

Linear probing of large language model (LLM) hidden states is widely used to claim that models learn distinct representations for different reasoning types. We test this by probing Qwen3-14B on three benchmarks spanning …