paper-with-me

홈 › Papers

What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA

2026-05-21 · Xinjie He, Zhiyuan Lin, Su Liu, Jialun Wu, Qiyang Xie, Weikai Zhou, Shuai Xiao arxiv

Reinforcement learning (RL) has emerged as a viable recipe for training LLM agents to reason over external memory banks in multi-session dialogue. Existing work trains exclusively on a single benchmark, leaving open how the composition of training data shapes the skills a memory agent acquires. We present a controlled empirical study that holds architecture, RL algorithm, and all hyperparameters fixed and varies only the training curriculum across three conditions: in-domain (LoCoMo), mixed-benchmark (LoCoMo + LongMemEval), and out-of-domain (LongMemEval only). Across two benchmarks and ten question types, curriculum composition acts as a fine-grained lever on specialization rather than a uniform scaling factor on performance. The mixed curriculum yields the strongest overall F1 on both evaluation sets. Training on a narrow out-of-domain set transfers a targeted skill - temporal reasoning - despite weak aggregate performance. Per-type differences substantially exceed aggregate differences, indicating that single-number benchmark comparisons systematically underreport curriculum effects. We further report two practical lessons from adapting GRPO to a single-GPU regime: cross-benchmark mixing requires filtering format-specific noise from memory banks to preserve training signal, and binary exact-match reward produces no learning signal at the small group sizes (G = 4) required on one GPU, motivating continuous reward functions in this regime.

📄 PDF Abstract BibTeX arXiv:2605.23067

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

When Memory Takes Gradients: Collaborative Vector Memory for Agentic Recommender Systems

2026-08-27 · Hanchong Chen, Xing Tang, Lingjie Li, Xiongfeng Shan 외 arxiv

Agentic recommender systems ground each decision of a large language model (LLM) in a persistent memory of the user, and in existing agents that memory is text: a narrative written and maintained by further LLM calls. Te…

Offspring from Reproduction Problems: What Replication Failure Teaches Us

2013-08-01 · ACL 2013 8 · Antske Fokkens, Marieke van Erp, Marten Postma, Ted Pedersen 외
Named Entity Recognition (NER)

Agentic Critical Training

2026-03-09 · Weize Liu, Minghui Liu, Sy-Tuyen Ho, Souradip Chakraborty 외 arxiv

Training large language models (LLMs) as autonomous agents often begins with imitation learning, but it only teaches agents what to do without understanding why: agents never contrast successful actions against suboptima…

Reinforcement LearningKnowledge Distillation

STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments

2026-05-28 · Junyang Wang, Haiyang Xu, Xi Zhang, Zhaoqing Zhu 외 arxiv

Mobile GUI agents excel at immediate reactive control but frequently fail in realistic, long-horizon tasks that require memory. This failure stems from a fundamental conflict between limited context windows and token-hea…

Reinforcement Learning

What Must Generalist Agents Remember?

2026-06-17 · Khurram Yamin, Namrata Deka, Maitreyi Swaroop, Albert Ting 외 arxiv

This paper develops a formal account of what generalist agents must store in memory in order to act near-optimally across multiple environments and goals. It shows that when two domains share an observational bottleneck …