paper-with-me

홈 › Papers

Leveraging weights signals -- Predicting and improving generalizability in reinforcement learning

2025-11-25 · Olivier Moulin, Vincent Francois-lavet, Paul Elbers, Mark Hoogendoorn arxiv

Generalizability of Reinforcement Learning (RL) agents (ability to perform on environments different from the ones they have been trained on) is a key problem as agents have the tendency to overfit to their training environments. In order to address this problem and offer a solution to increase the generalizability of RL agents, we introduce a new methodology to predict the generalizability score of RL agents based on the internal weights of the agent's neural networks. Using this prediction capability, we propose some changes in the Proximal Policy Optimization (PPO) loss function to boost the generalization score of the agents trained with this upgraded version. Experimental results demonstrate that our improved PPO algorithm yields agents with stronger generalizability compared to the original version.

📄 PDF Abstract BibTeX arXiv:2511.20234

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Learning to Learn Weight Generation via Local Consistency Diffusion

2025-02-03 · Yunchuan Guan, Yu Liu, Ke Zhou, Zhiqi Shen 외

Diffusion-based algorithms have emerged as promising techniques for weight generation. However, existing solutions are limited by two challenges: generalizability and local target assignment. The former arises from the i…

Domain GeneralizationFew-Shot LearningLanguage ModelingLanguage Modelling+4

Predicting Heart Activity from Speech using Data-driven and Knowledge-based features

2024-06-10 · Gasser Elbanna, Zohreh Mostaani, Mathew Magimai. -Doss

Accurately predicting heart activity and other biological signals is crucial for diagnosis and monitoring. Given that speech is an outcome of multiple physiological systems, a significant body of work studied the acousti…

Reinforcement Learning for Optimal Experiment Design in Parameter Identification of Mechatronic Systems

2026-05-19 · Julian Langschwert, Georg Schaefer, Jakob Rehrl, Stefan Huber 외 arxiv

Informative excitation signals are critical for accurate system identification of mechatronic systems, yet classical system identification (SI) approaches require expert knowledge and hand-crafted signal design to respec…

Reinforcement Learning

Reinforcement Logic Rule Learning for Temporal Point Processes

2023-08-11 · Chao Yang, Lu Wang, Kun Gao, Shuang Li

We propose a framework that can incrementally expand the explanatory temporal logic rule set to explain the occurrence of temporal events. Leveraging the temporal point process modeling and learning framework, the rule c…

Point Processes

R$^3$L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification

2026-01-07 · Weijie Shi, Yanxi Chen, Zexi Li, Xuchen Pan 외 arxiv

Reinforcement learning drives recent advances in LLM reasoning and agentic capabilities, yet current approaches struggle with both exploration and exploitation. Exploration suffers from low success rates on difficult tas…

Reinforcement Learning