paper-with-me

홈 › Papers

Explainable RL Policies by Distilling to Locally-Specialized Linear Policies with Voronoi State Partitioning

2025-11-17 · Senne Deproost, Dennis Steckelmacher, Ann Nowé arxiv

Deep Reinforcement Learning is one of the state-of-the-art methods for producing near-optimal system controllers. However, deep RL algorithms train a deep neural network, that lacks transparency, which poses challenges when the controller has to meet regulations, or foster trust. To alleviate this, one could transfer the learned behaviour into a model that is human-readable by design using knowledge distilla- tion. Often this is done with a single model which mimics the original model on average but could struggle in more dynamic situations. A key challenge is that this simpler model should have the right balance be- tween flexibility and complexity or right balance between balance bias and accuracy. We propose a new model-agnostic method to divide the state space into regions where a simplified, human-understandable model can operate in. In this paper, we use Voronoi partitioning to find regions where linear models can achieve similar performance to the original con- troller. We evaluate our approach on a gridworld environment and a classic control task. We observe that our proposed distillation to locally- specialized linear models produces policies that are explainable and show that the distillation matches or even slightly outperforms the black-box policy they are distilled from.

📄 PDF Abstract BibTeX arXiv:2511.13322

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models

2026-05-14 · Senne Deproost, Denis Steckelmacher, Ann Nowé arxiv

Despite many successful attempts at explaining Deep Reinforcement Learning policies using distillation, it remains difficult to balance the performance-interpretability trade-off and select a fitting surrogate model. In …

Reinforcement Learning

Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability

2017-03-17 · ICML 2017 8 · Shayegan Omidshafiei, Jason Pazis, Christopher Amato, Jonathan P. How 외

Many real-world tasks involve multiple agents with partial observability and limited communication. Learning is challenging in these settings due to local viewpoints of agents, which perceive the world as non-stationary …

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Locally Linear Continual Learning for Time Series based on VC-Theoretical Generalization Bounds

2026-03-14 · Yan V. G. Ferreira, Igor B. Lima, Pedro H. G. Mapa S., Felipe V. Campos 외 arxiv

Most machine learning methods assume fixed probability distributions, limiting their applicability in nonstationary real-world scenarios. While continual learning methods address this issue, current approaches often rely…

Time Series ForecastingContinual Learning

Explainable LLM-driven Multi-dimensional Distillation for E-Commerce Relevance Learning

2024-11-20 · Gang Zhao, XiMing Zhang, Chenji Lu, Hui Zhao 외

Effective query-item relevance modeling is pivotal for enhancing user experience and safeguarding user satisfaction in e-commerce search systems. Recently, benefiting from the vast inherent knowledge, Large Language Mode…

Knowledge DistillationLarge Language Model

Distilling Reinforcement Learning Policies for Interpretable Robot Locomotion: Gradient Boosting Machines and Symbolic Regression

2024-03-21 · Fernando Acero, Zhibin Li

Recent advancements in reinforcement learning (RL) have led to remarkable achievements in robot locomotion capabilities. However, the complexity and ``black-box'' nature of neural network-based RL policies hinder their i…

Additive modelsReinforcement Learning (RL)Symbolic Regression