paper-with-me

홈 › Papers

More Than One Teacher: Adaptive Multi-Guidance Policy Optimization for Diverse Exploration

2025-10-02 · Xiaoyang Yuan, Yujuan Ding, Yi Bin, Wenqi Shao, Jinyu Cai, Jingkuan Song, Yang Yang, Heng Tao Shen arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) is a promising paradigm for enhancing the reasoning ability in Large Language Models (LLMs). However, prevailing methods primarily rely on self-exploration or a single off-policy teacher to elicit long chain-of-thought (LongCoT) reasoning, which may introduce intrinsic model biases and restrict exploration, ultimately limiting reasoning diversity and performance. Drawing inspiration from multi-teacher strategies in knowledge distillation, we introduce Adaptive Multi-Guidance Policy Optimization (AMPO), a novel framework that adaptively leverages guidance from multiple proficient teacher models, but only when the on-policy model fails to generate correct solutions. This "guidance-on-demand" approach expands exploration while preserving the value of self-discovery. Moreover, AMPO incorporates a comprehension-based selection mechanism, prompting the student to learn from the reasoning paths that it is most likely to comprehend, thus balancing broad exploration with effective exploitation. Extensive experiments show AMPO substantially outperforms a strong baseline (GRPO), with a 4.3% improvement on mathematical reasoning tasks and 12.2% on out-of-distribution tasks, while significantly boosting Pass@k performance and enabling more diverse exploration. Notably, using four peer-sized teachers, our method achieves comparable results to approaches that leverage a single, more powerful teacher (e.g., DeepSeek-R1) with more data. These results demonstrate a more efficient and scalable path to superior reasoning and generalizability. Our code is available at https://github.com/SII-Enigma/AMPO.

📄 PDF Abstract BibTeX arXiv:2510.02227

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningKnowledge DistillationMathematical Reasoning

Similar Papers 제목 키워드 기반

Point Adversarial Self Mining: A Simple Method for Facial Expression Recognition

2020-08-26 · Ping Liu, Yuewei Lin, Zibo Meng, Lu Lu 외

In this paper, we propose a simple yet effective approach, named Point Adversarial Self Mining (PASM), to improve the recognition accuracy in facial expression recognition. Unlike previous works focusing on designing spe…

Adversarial AttackData AugmentationFacial Expression RecognitionFacial Expression Recognition (FER)+2

Safety-Regulated Transfer Reinforcement Learning with Adaptive Teacher Guidance

2026-06-25 · Wenjie Huang, Yang Li, Jingjia Teng, Mingwei Jin 외 arxiv

We propose Safety-Regulated Adaptive Transfer Reinforcement Learning (SRATRL), a teacher--student framework that combines safety-triggered intervention, safety-adaptive value shaping, and policy-compatibility-based optim…

Reinforcement LearningTransfer LearningDomain Adaptation

Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models

2026-03-22 · Jingchen Sun, Shaobo Han, Deep Patel, Wataru Kohno 외 arxiv

Knowledge distillation establishes a learning paradigm that leverages both data supervision and teacher guidance. However, determining the optimal balance between learning from data and learning from the teacher is chall…

Knowledge Distillation

AdaGAT: Adaptive Guidance Adversarial Training for the Robustness of Deep Neural Networks

2025-08-24 · Zhenyu Liu, Huizhi Liang, Xinrun Li, Vaclav Snasel 외 arxiv

Adversarial distillation (AD) is a knowledge distillation technique that facilitates the transfer of robustness from teacher deep neural network (DNN) models to lightweight target (student) DNN models, enabling the targe…

Knowledge Distillation

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance

2026-08-01 · Zhuowen Han, Jinwei Xiao, Zhengxi Lu, Renren Jin 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO) is widely adopted, it suffers from spar…

Reinforcement Learning