paper-with-me

홈 › Papers

Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities

2025-08-27 · Rikuto Kotoge, Mai Nishimura, Jiaxin Ma arxiv

Reinforcement Learning has emerged as a dominant post-training approach to elicit agentic RAG behaviors such as search and planning from language models. Despite its success with larger models, applying RL to compact models (e.g., 0.5--1B parameters) presents unique challenges. The compact models exhibit poor initial performance, resulting in sparse rewards and unstable training. To overcome these difficulties, we propose Distillation-Guided Policy Optimization (DGPO), which employs cold-start initialization from teacher demonstrations and continuous teacher guidance during policy optimization. To understand how compact models preserve agentic behavior, we introduce Agentic RAG Capabilities (ARC), a fine-grained metric analyzing reasoning, search coordination, and response synthesis. Comprehensive experiments demonstrate that DGPO enables compact models to achieve sophisticated agentic search behaviors, even outperforming the larger teacher model in some cases. DGPO makes agentic RAG feasible in computing resource-constrained environments.

📄 PDF Abstract BibTeX arXiv:2508.20324

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Structured Agent Distillation for Large Language Model

2025-05-20 · Jun Liu, Zhenglun Kong, Peiyan Dong, Changdi Yang 외

Large language models (LLMs) exhibit strong capabilities as decision-making agents by interleaving reasoning and actions, as seen in ReAct-style frameworks. Yet, their practical deployment is constrained by high inferenc…

Decision MakingImitation LearningLanguage ModelingLanguage Modelling+2

Scalable Behaviour Cloning on Browser Using via Skill Distillation

2026-06-30 · Kaisen Yang, Zheng Jiang, Yuzhao Peng, Houde Qian 외 arxiv

Internet users collectively perform an enormous range of skilled work through web browsers, from software development and document editing to search, forms, and enterprise workflows, making human browsing a highly scalab…

Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents

2026-06-10 · Kushal Raj Bhandari, Ling Yue, Ching-Yun Ko, Dhaval Patel 외 arxiv

Compact language models (LMs) reduce cost, latency, and deployment risk for tool agents. Yet MCP-style tool use requires more than isolated function calling: an agent must discover tools from live catalogs, satisfy schem…

Structured Pruning Learns Compact and Accurate Models

2021-11-16 · ACL ARR November 2021 11 · Anonymous

The growing size of neural language models has led to increased attention in model compression. Pruning methods start from a large model and gradually remove model weights---they can significantly reduce the model size b…

Model Compression

DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation

2026-06-29 · Peyman Hosseini, Ondrej Bohdal, Ahmed Alajrami, Andrea Maracani 외 hf

Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over multiple turns, but this ability typically depends on large models, long contexts, and repeated inference c…