paper-with-me

홈 › Papers

TRACED: Transition-aware Regret Approximation with Co-learnability for Environment Design

2025-06-24 · Geonwoo Cho, Jaegyun Im, JIhwan Lee, Hojun Yi, Sejin Kim, Sundong Kim

Generalizing deep reinforcement learning agents to unseen environments remains a significant challenge. One promising solution is Unsupervised Environment Design (UED), a co-evolutionary framework in which a teacher adaptively generates tasks with high learning potential, while a student learns a robust policy from this evolving curriculum. Existing UED methods typically measure learning potential via regret, the gap between optimal and current performance, approximated solely by value-function loss. Building on these approaches, we introduce the transition prediction error as an additional term in our regret approximation. To capture how training on one task affects performance on others, we further propose a lightweight metric called co-learnability. By combining these two measures, we present Transition-aware Regret Approximation with Co-learnability for Environment Design (TRACED). Empirical evaluations show that TRACED yields curricula that improve zero-shot generalization across multiple benchmarks while requiring up to 2x fewer environment interactions than strong baselines. Ablation studies confirm that the transition prediction error drives rapid complexity ramp-up and that co-learnability delivers additional gains when paired with the transition prediction error. These results demonstrate how refined regret approximation and explicit modeling of task relationships can be leveraged for sample-efficient curriculum design in UED.

📄 PDF Abstract BibTeX arXiv:2506.19997

Code (1)

cho-geonwoo/traced 공식 구현 pytorch

Tasks

Deep Reinforcement LearningZero-shot Generalization

Similar Papers 제목 키워드 기반

No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery

2024-08-27 · Alexander Rutherford, Michael Beukman, Timon Willi, Bruno Lacerda 외

What data or environments to use for training to improve downstream performance is a longstanding and very topical question in reinforcement learning. In particular, Unsupervised Environment Design (UED) methods have gai…

Balancing Expressivity and Learnability in Quantum Kernel Bandit Optimization

2026-07-01 · Yuqi Huang, Vincent Y. F. Tan, Sharu Theresa Jose arxiv

We investigate Gaussian process (GP) bandit optimization with quantum kernels, assuming the mean reward function lies in the reproducing kernel Hilbert space (RKHS) induced by the quantum kernel. This setting is motivate…

Value Function Approximations via Kernel Embeddings for No-Regret Reinforcement Learning

2020-11-16 · Sayak Ray Chowdhury, Rafael Oliveira

We consider the regret minimization problem in reinforcement learning (RL) in the episodic setting. In many real-world RL environments, the state and action spaces are continuous or very large. Existing approaches establ…

reinforcement-learningReinforcement Learning (RL)

Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation

2024-05-30 · Wooseong Cho, TaeHyun Hwang, Joongkyu Lee, Min-hwan Oh

We study reinforcement learning with multinomial logistic (MNL) function approximation where the underlying transition probability kernel of the Markov decision processes (MDPs) is parametrized by an unknown transition c…

reinforcement-learningReinforcement Learning

From Confounding to Learning: Dynamic Service Fee Pricing on Third-Party Platforms

2025-12-28 · Rui Ai, David Simchi-Levi, Feng Zhu arxiv

We study the pricing behavior of third-party platforms facing strategic agents. Assuming the platform is a revenue maximizer, it observes market features that generally affect demand. Since only transacted quantities and…