CHIRPs: Change-Induced Regret Proxy metrics for Lifelong Reinforcement Learning
Reinforcement learning (RL) agents are costly to train and fragile to environmental changes. They often perform poorly when there are many changing tasks, prohibiting their widespread deployment in the real world. Many Lifelong RL agent designs have been proposed to mitigate issues such as catastrophic forgetting or demonstrate positive characteristics like forward transfer when change occurs. However, no prior work has established whether the impact on agent performance can be predicted from the change itself. Understanding this relationship will help agents proactively mitigate a change's impact for improved learning performance. We propose Change-Induced Regret Proxy (CHIRP) metrics to link change to agent performance drops and use two environments to demonstrate a CHIRP's utility in lifelong learning. A simple CHIRP-based agent achieved $48\%$ higher performance than the next best method in one benchmark and attained the best success rates in 8 of 10 tasks in a second benchmark which proved difficult for existing lifelong RL agents.
Code (0)
등록된 구현이 없습니다.
Tasks
Lifelong learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Automated pulse discrimination of two freely-swimming weakly electric fish and analysis of their electrical behavior during a dominance contest
Electric fishes modulate their electric organ discharges with a remarkable variability. Some patterns can be easily identified, such as pulse rate changes, offs and chirps, which are often associated with important behav…
PACE: Parameter Change for Unsupervised Environment Design
Unsupervised Environment Design (UED) offers a promising paradigm for improving reinforcement learning generalization by adaptively shaping training environments, but it requires reliable environment evaluation to remain…
Reinforcement LearningDecision-Weighted Flow Matching for Contextual Stochastic Optimization
Conditional generative models are increasingly used as scenario generators for stochastic optimization, but standard training objectives emphasize uniform distributional fit rather than the downstream decisions induced b…
Stochastic OptimizationOn the Convergence of No-Regret Dynamics in Information Retrieval Games with Proportional Ranking Functions
Publishers who publish their content on the web act strategically, in a behavior that can be modeled within the online learning framework. Regret, a central concept in machine learning, serves as a canonical measure for …
Information RetrievalRetrievalWideband Index Modulation with Circularly-Shifted Chirps
In this study, we propose a wideband index modulation (IM) based on circularly-shifted chirps. To derive the proposed method, we first prove that a Golay complementary pair (GCP) can be constructed by linearly combining …