A Survey on Explainable Deep Reinforcement Learning
Deep Reinforcement Learning (DRL) has achieved remarkable success in sequential decision-making tasks across diverse domains, yet its reliance on black-box neural architectures hinders interpretability, trust, and deployment in high-stakes applications. Explainable Deep Reinforcement Learning (XRL) addresses these challenges by enhancing transparency through feature-level, state-level, dataset-level, and model-level explanation techniques. This survey provides a comprehensive review of XRL methods, evaluates their qualitative and quantitative assessment frameworks, and explores their role in policy refinement, adversarial robustness, and security. Additionally, we examine the integration of reinforcement learning with Large Language Models (LLMs), particularly through Reinforcement Learning from Human Feedback (RLHF), which optimizes AI alignment with human preferences. We conclude by highlighting open research challenges and future directions to advance the development of interpretable, reliable, and accountable DRL systems.
Code (0)
등록된 구현이 없습니다.
Tasks
Adversarial RobustnessDecision MakingDeep Reinforcement Learningreinforcement-learningReinforcement LearningSequential Decision MakingSurveySimilar Papers 제목 키워드 기반
A Survey of Explainable Reinforcement Learning
Explainable reinforcement learning (XRL) is an emerging subfield of explainable machine learning that has attracted considerable attention in recent years. The goal of XRL is to elucidate the decision-making process of l…
Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+2A Survey on Explainable Reinforcement Learning: Concepts, Algorithms, Challenges
Reinforcement Learning (RL) is a popular machine learning paradigm where intelligent agents interact with the environment to fulfill a long-term goal. Driven by the resurgence of deep learning, Deep RL (DRL) has witnesse…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)SurveyA Survey of Explainable Reinforcement Learning: Targets, Methods and Needs
The success of recent Artificial Intelligence (AI) models has been accompanied by the opacity of their internal mechanisms, due notably to the use of deep neural networks. In order to understand these internal mechanisms…
reinforcement-learningReinforcement LearningExplainable Reinforcement Learning: A Survey
Explainable Artificial Intelligence (XAI), i.e., the development of more transparent and interpretable AI models, has gained increased traction over the last few years. This is due to the fact that, in conjunction with t…
Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)Philosophyreinforcement-learning+3Explainable Recommendation: A Survey and New Perspectives
Explainable recommendation attempts to develop models that generate not only high-quality recommendations but also intuitive explanations. The explanations may either be post-hoc or directly come from an explainable mode…
Explainable RecommendationPersuasivenessProduct RecommendationRecommendation Systems+1