paper-with-me

홈 › Papers

Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration

2026-01-12 · Yang Zhao, Yangou Ouyang, Xiao Ding, Hepeng Wang, Bibo Cai, Kai Xiong, Jinglong Gao, Zhouhao Sun, Li Du, Bing Qin, Ting Liu arxiv

While Hybrid Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become the standard paradigm for training LLM agents, effective mechanisms for data allocation between these stages remain largely underexplored. Current data arbitration strategies often rely on surface-level heuristics that fail to diagnose intrinsic learning needs. Since SFT targets pattern consolidation through imitation while RL drives structural adaptation via exploration, misaligning data with these functional roles causes severe optimization interference. We propose PRISM, a dynamics-aware framework grounded in Schema Theory that arbitrates data based on its degree of cognitive conflict with the model's existing knowledge. By analyzing the spatial geometric structure of gradients, PRISM identifies data triggering high spatial concentration as high-conflict signals that require RL for structural restructuring. In contrast, data yielding diffuse updates is routed to SFT for efficient consolidation. Extensive experiments on WebShop and ALFWorld demonstrate that PRISM achieves a Pareto improvement, outperforming state-of-the-art hybrid methods while reducing computational costs by up to 3.22$\times$. Our findings suggest that disentangling data based on internal optimization regimes is crucial for scalable and robust agent alignment.

📄 PDF Abstract BibTeX arXiv:2601.07224

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs

2025-10-22 · Xinfeng Xia, Jiacheng Liu, Xiaofeng Hou, Peng Tang 외 arxiv

Mixture-of-Experts (MoE) models, the state-of-the-art in large-scale AI, achieve high quality by sparsely activating parameters. However, their reliance on routing between a few monolithic experts via a top-k mechanism c…

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

2026-08-15 · Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi hf

Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signa…

Test-time Adaptation

Gated Differentiable Working Memory for Long-Context Language Modeling

2026-01-19 · Lingrui Mei, Shenghua Liu, Yiwei Wang, Yuyao Ge 외 arxiv

Long contexts challenge transformers: attention scores dilute across thousands of tokens, critical information is often lost in the middle, and models struggle to adapt to novel patterns at inference time. Recent work on…

Test-time Adaptation

PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection

2025-10-27 · Yusu Qian, Cheng Wan, Chao Jia, Yinfei Yang 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable progress on vision-language tasks, yet their reasoning processes remain sometimes unreliable. We introduce PRISM-Bench, a benchmark of puzzle-based visual…

Multimodal ReasoningAnswer GenerationVisual Reasoning

Prism: An Evolutionary Memory Substrate for Multi-Agent Open-Ended Discovery

2026-04-08 · Suyash Mishra arxiv

We introduce \prism{} (\textbf{P}robabilistic \textbf{R}etrieval with \textbf{I}nformation-\textbf{S}tratified \textbf{M}emory), an evolutionary memory substrate for multi-agent AI systems engaged in open-ended discovery…

Information Retrieval