paper-with-me

Papers

MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement

2025-08-13 · Weitao Jia, Jinghui Lu, Haiyang Yu, Siqi Wang, Guozhi Tang, An-Lan Wang, Weijie Yin, Dingkang Yang, Yuxiang Nie, Bin Shan, Hao Feng, Irene Li, Kun Yang, Han Wang, Jingqun Tang, Teng Fu, Changhong Jin, Chao Feng, Xiaohui Lv, Can Huang arxiv

Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers provide no learning signal, particularly in challenging tasks. To address this, we propose Multi-Expert Mutual Learning GRPO (MEML-GRPO), an innovative framework that utilizes diverse expert prompts as system prompts to generate a broader range of responses, substantially increasing the likelihood of identifying correct solutions. Additionally, we introduce an inter-expert mutual learning mechanism that facilitates knowledge sharing and transfer among experts, further boosting the model's performance through RLVR. Extensive experiments across multiple reasoning benchmarks show that MEML-GRPO delivers significant improvements, achieving an average performance gain of 4.89% with Qwen and 11.33% with Llama, effectively overcoming the core limitations of traditional RLVR methods.

📄 PDF Abstract BibTeX arXiv:2508.09670

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models

2026-05-08 · Xiaoze Liu, Dhananjay Ram, Yuting Zhang, Zhaoyang Zhang 외 arxiv

We introduce Mutual Reinforcement Learning, a framework for concurrent RL post-training in which heterogeneous LLM policies exchange typed experience while keeping separate parameters, objectives, and tokenizers. The fra…

Reinforcement Learning

MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents

2026-07-28 · Shuyue Wei, Chang Liu, Zimu Zhou, Yongxin Tong 외 arxiv

Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon reasoning, personalized responses, and knowledge reuse. However, existing LLM memory systems typically…

MemLoRA: Distilling Expert Adapters for On-Device Memory Systems

2025-12-04 · Massimo Bini, Ondrej Bohdal, Umberto Michieli, Zeynep Akata 외 arxiv

Memory-augmented Large Language Models (LLMs) have demonstrated remarkable consistency during prolonged dialogues by storing relevant memories and incorporating them as context. Such memory-based personalization is also …

Visual Question AnsweringKnowledge DistillationVisual Reasoning

TimeML-strict: clarifying temporal annotation

2013-04-26 · Leon Derczynski, Hector Llorens, Naushad UzZaman

TimeML is an XML-based schema for annotating temporal information over discourse. The standard has been used to annotate a variety of resources and is followed by a number of tools, the creation of which constitute hundr…

valid

MemLong: Memory-Augmented Retrieval for Long Text Modeling

2024-08-30 · Weijie Liu, Zecheng Tang, Juntao Li, Kehai Chen 외

Recent advancements in Large Language Models (LLMs) have yielded remarkable success across diverse fields. However, handling long contexts remains a significant challenge for LLMs due to the quadratic time and space comp…

4kDecoderGPUInformation Retrieval+4