paper-with-me

홈 › Papers

Multiple Token Divergence: Measuring and Steering In-Context Computation Density

2025-12-28 · Vincent Herrmann, Eric Alcaide, Michael Wand, Jürgen Schmidhuber arxiv

Measuring the in-context computational effort of language models is a key challenge, as metrics like next-token loss fail to capture reasoning complexity. Prior methods based on latent state compressibility can be invasive and unstable. We propose Multiple Token Divergence (MTD), a simple measure of computational effort defined as the KL divergence between a model's full output distribution and that of a shallow, auxiliary prediction head. MTD can be computed directly from pre-trained models with multiple prediction heads, requiring no additional training. Building on this, we introduce Divergence Steering, a novel decoding method to control the computational character of generated text. We empirically show that MTD is more effective than prior methods at distinguishing complex tasks from simple ones. On mathematical reasoning benchmarks, MTD correlates positively with problem difficulty. Lower MTD is associated with more accurate reasoning. MTD provides a practical, lightweight tool for analyzing and steering the computational dynamics of language models.

📄 PDF Abstract BibTeX arXiv:2512.22944

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Don't Lose Focus: Activation Steering via Key-Orthogonal Projections

2026-05-07 · Haoyan Luo, Mateo Espinosa Zarlenga, Mateja Jamnik arxiv

Activation steering controls LLM behaviour towards target behaviour by intervening in internal representations, yet it often degrades reasoning and retrieval performance. We argue that a primary cause of this trade-off i…

Compositional Steering of Large Language Models with Steering Tokens

2026-01-08 · Gorjan Radevski, Kiril Gashteovski, Giwon Hong, Carolin Lawrence 외 arxiv

Deploying LLMs in real-world applications requires controllable output that satisfies multiple desiderata at the same time. While existing work extensively addresses LLM steering for a single behavior, \textit{compositio…

Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning

2026-04-07 · Yanbei Jiang, Amr Keleg, Ryandito Diandaru, Jey Han Lau 외 arxiv

While the real world is inherently stochastic, Large Language Models (LLMs) are predominantly evaluated on single-round inference against fixed ground truths. In this work, we shift the lens to distribution alignment: as…

Prompt Engineering

Language steering in latent space to mitigate unintended code-switching

2025-10-11 · Andrey Goncharov, Nikolai Kondusov, Alexey Zaytsev arxiv

Multilingual Large Language Models (LLMs) often exhibit hallucinations such as unintended code-switching, reducing reliability in downstream tasks. We propose latent-space language steering, a lightweight inference-time …

Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders

2025-10-04 · Xu Wang, Yan Hu, Benyou Wang, Difan Zou arxiv

Sparse Autoencoders (SAEs) are widely used to steer large language models (LLMs), based on the assumption that their interpretable features naturally enable effective model behavior steering. Yet, a fundamental question …