paper-with-me

홈 › Papers

Language steering in latent space to mitigate unintended code-switching

2025-10-11 · Andrey Goncharov, Nikolai Kondusov, Alexey Zaytsev arxiv

Multilingual Large Language Models (LLMs) often exhibit hallucinations such as unintended code-switching, reducing reliability in downstream tasks. We propose latent-space language steering, a lightweight inference-time method that identifies language directions via Principal Component Analysis (PCA) on parallel translations and steers token embeddings along these axes to control language identity. Our approach mitigates code-switching while preserving semantics with negligible computational overhead and requires only minimal parallel data for calibration. Empirically, we achieve 95-99\% language classification accuracy using a single principal component and reduce next-token distributional divergence by up to 55\% across multiple language pairs on Qwen2.5 and Llama-3.2 models. Generation-based evaluation on Llama-3.2 further demonstrates 63--99\% reduction in Code-Switching Index across four language pairs ($p < 0.001$). We further analyze the layer-wise evolution of language representations, revealing that language identity concentrates in final layers with near-perfect linear separability.

📄 PDF Abstract BibTeX arXiv:2510.13849

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

2026-07-22 · Kavin Aravindan, Arihant Rastogi, Aadi Prasad, Krishak Aneja 외 arxiv

Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refu…

Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning

2025-07-22 · Helena Casademunt, Caden Juang, Adam Karvonen, Samuel Marks 외 arxiv

Fine-tuning large language models (LLMs) can lead to unintended out-of-distribution generalization. Standard approaches to this problem rely on modifying training data, for example by adding data that better specify the …

Steering Language Models with Weight Arithmetic

2025-11-07 · Constanza Fierro, Fabien Roger arxiv

Providing high-quality feedback to Large Language Models (LLMs) on a diverse training distribution can be difficult and expensive, and providing feedback only on a narrow distribution can result in unintended generalizat…

Improving Steering Vectors by Targeting Sparse Autoencoder Features

2024-11-04 · Sviatoslav Chalnev, Matthew Siu, Arthur Conmy

To control the behavior of language models, steering methods attempt to ensure that outputs of the model satisfy specific pre-defined properties. Adding steering vectors to the model is a promising method of model contro…

Contrastive Conceptor Activation Steering (COAST): Unlocking Vision-Language-Action Models through Hidden States

2026-05-16 · Miranda Muqing Miao, Subin Kim, Brandon Yang, Lyle Ungar arxiv

Vision-Language-Action (VLA) models leverage powerful perceptual priors from web-scale Vision-Language Model (VLM) pre-training, yet they remain surprisingly brittle in practice, frequently failing at simple robotic task…