paper-with-me

홈 › Papers

Adaptive Teaching in Heterogeneous Agents: Balancing Surprise in Sparse Reward Scenarios

2024-05-23 · Emma Clark, Kanghyun Ryu, Negar Mehr

Learning from Demonstration (LfD) can be an efficient way to train systems with analogous agents by enabling `Student'' agents to learn from the demonstrations of the most experienced Teacher'' agent, instead of training their policy in parallel. However, when there are discrepancies in agent capabilities, such as divergent actuator power or joint angle constraints, naively replicating demonstrations that are out of bounds for the Student's capability can limit efficient learning. We present a Teacher-Student learning framework specifically tailored to address the challenge of heterogeneity between the Teacher and Student agents. Our framework is based on the concept of `surprise'', inspired by its application in exploration incentivization in sparse-reward environments. Surprise is repurposed to enable the Teacher to detect and adapt to differences between itself and the Student. By focusing on maximizing its surprise in response to the environment while concurrently minimizing the Student's surprise in response to the demonstrations, the Teacher agent can effectively tailor its demonstrations to the Student's specific capabilities and constraints. We validate our method by demonstrating improvements in the Student's learning in control tasks within sparse-reward environments.

📄 PDF Abstract BibTeX arXiv:2405.14199

Code (1)

labicon/Surprise_based_Teaching 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Class Teaching for Inverse Reinforcement Learners

2019-11-29 · Manuel Lopes, Francisco Melo

In this paper we propose the first machine teaching algorithm for multiple inverse reinforcement learners. Specifically, our contributions are: (i) we formally introduce the problem of teaching a sequential task to a het…

SMiRL: Surprise Minimizing RL in Entropic Environments

2019-09-25 · Glen Berseth, Daniel Geng, Coline Devin, Dinesh Jayaraman 외

All living organisms struggle against the forces of nature to carve out niches where they can maintain relative stasis. We propose that such a search for order amidst chaos might offer a unifying principle for the emerge…

Unsupervised Pre-trainingUnsupervised Reinforcement Learning

Heterogeneous Trader Responses to Macroeconomic Surprises: Simulating Order Flow Dynamics

2025-05-04 · Haochuan Wang

Understanding how market participants react to shocks like scheduled macroeconomic news is crucial for both traders and policymakers. We develop a calibrated data generation process DGP that embeds four stylized trader a…

Investigating Pedagogical Teacher and Student LLM Agents: Genetic Adaptation Meets Retrieval Augmented Generation Across Learning Style

2025-05-25 · Debdeep Sanyal, Agniva Maiti, Umakanta Maharana, Dhruv Kumar 외

Effective teaching requires adapting instructional strategies to accommodate the diverse cognitive and behavioral profiles of students, a persistent challenge in education and teacher training. While Large Language Model…

RAGRetrievalRetrieval-augmented Generation

LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching

2026-06-15 · Jaward Sesay, Yue Yu, Siwei Dong, Börje F. Karlsson arxiv

Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but also dynamically adapt their instruction to diverse learners. However, existing …

Semantic Segmentation