paper-with-me

Papers

Steering Neural Network Training through Interpretable Constraints Based on Partial Dependence

2026-07-09 · Yann Claes, Pierre Geurts, Vân Anh Huynh-Thu arxiv

Over the last few years, there has been an increased interest in making machine learning models more interpretable. Although a great deal of effort goes into developing techniques for interpreting the interactions learned by a given model, fewer studies focus on assessing the quality of such explanations. Even fewer focus on how to adjust the model to produce explanations faithful to prior knowledge, a process known as explanation-guided learning. Furthermore, most approaches in this area focus on classification problems and usually assume prior knowledge about which input features or regions are most important. In this work, we introduce a new approach to steering neural networks based on partial dependence, such that their average response to certain features aligns with specific functional domain knowledge about the problem. We empirically demonstrate on a range of regression problems, including dynamical systems forecasting, that models whose training has been controlled using our method perform better than unconstrained models and are more data-efficient. Moreover, we highlight that interpretations obtained from the former actually align with the user-provided knowledge, whereas those obtained from the latter do not.

📄 PDF Abstract BibTeX arXiv:2607.08641

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM-Augmented Semantic Steering of Text Embedding Projection Spaces

2026-05-03 · Wei Liu, Eric Krokos, Kirsten Whitley, Rebecca Faust 외 arxiv

Low-dimensional projections of text embeddings support visual analysis of document collections, but their spatial organization may not reflect the relationships an analyst intends to examine. Existing semantic interactio…

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

2026-04-09 · Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha arxiv

Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explanation for how it works-- specifically, what internal mechanisms steeri…

The Rogue Scalpel: Activation Steering Compromises LLM Safety

2025-09-26 · Anton Korznikov, Andrey Galichin, Alexey Dontsov, Oleg Y. Rogov 외 arxiv

Activation steering is a promising technique for controlling LLM behavior by adding semantically meaningful vectors directly into a model's hidden states during inference. It is often framed as a precise, interpretable, …

Steer2Edit: From Activation Steering to Component-Level Editing

2026-02-10 · Chung-En Sun, Ge Yan, Zimo Wang, Tsui-Wei Weng arxiv

Steering methods influence Large Language Model behavior by identifying semantic directions in hidden representations, but are typically realized through inference-time activation interventions that apply a fixed, global…

Decision-Driven Geosteering Under Uncertainty: A Unified Framework for Sequential Decision Optimization

2026-06-15 · Hibat Errahmen Djecta, Sergey Alyaev, Kristian Fossum, Reidar B. Bratvold 외 arxiv

Geosteering requires navigating a well trajectory through an unknown geological configuration, while sequentially updating decisions based on indirect measurements acquired during drilling. This work presents an uncertai…

Reinforcement Learning