paper-with-me

Papers

AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations

2026-03-02 · Cheng Jiayang, Dongyu Ru, Lin Qiu, Yiyang Li, Xuezhi Cao, Yangqiu Song, Xunliang Cai arxiv

Long-horizon interactions between users and LLM-based assistants necessitate effective memory management, yet current approaches face challenges in training and evaluation of memory. Existing memory benchmarks rely on static, off-policy data as context, limiting evaluation reliability and scalability. To address these gaps, we introduce AMemGym, an interactive environment enabling on-policy evaluation and optimization for memory-driven personalization. AMemGym employs structured data sampling to predefine user profiles, state-dependent questions, and state evolution trajectories, enabling cost-effective generation of high-quality, evaluation-aligned interactions. LLM-simulated users expose latent states through role-play while maintaining structured state consistency. Comprehensive metrics based on structured data guide both assessment and optimization of assistants. Extensive experiments reveal performance gaps in existing memory systems (e.g., RAG, long-context LLMs, and agentic memory) and corresponding reasons. AMemGym not only enables effective selection among competing approaches but also can potentially drive the self-evolution of memory management strategies. By bridging structured state evolution with free-form interactions, our framework provides a scalable, diagnostically rich environment for advancing memory capabilities in conversational agents.

📄 PDF Abstract BibTeX arXiv:2603.01966

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory

2024-10-14 · Di wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang 외

Recent large language model (LLM)-driven chat assistant systems have integrated memory components to track user-assistant chat histories, enabling more accurate and personalized responses. However, their long-term memory…

BenchmarkingLarge Language ModelQuestion Answering

TeleEgo: Benchmarking Egocentric AI Assistants in the Wild

2025-10-28 · Jiaqi Yan, Ruilong Ren, Jingren Liu, Shuning Xu 외 arxiv

Egocentric AI assistants in real-world settings must process multi-modal inputs (video, audio, text), respond in real time, and retain evolving long-term memory. However, existing benchmarks typically evaluate these abil…

HiMeS: Hippocampus-inspired Memory System for Personalized AI Assistants

2026-01-06 · Hailong Li, Feifei Li, Wenhui Que, Xingyu Fan arxiv

Large language models (LLMs) power many interactive systems such as chatbots, customer-service agents, and personal assistants. In knowledge-intensive scenarios requiring user-specific personalization, conventional retri…

Reinforcement Learning

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory

2025-11-25 · Tianxin Wei, Noveen Sachdeva, Benjamin Coleman, Zhankui He 외 arxiv

Statefulness is essential for large language model (LLM) agents to perform long-term planning and problem-solving. This makes memory a critical component, yet its management and evolution remain largely underexplored. Ex…

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory

2026-05-30 · Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, James Fort 외 arxiv

AI glasses present a compelling platform for AI agents to serve as personalized memory assistants. To be genuinely useful, such systems must move beyond short-term video comprehension and address memory gaps that humans …

Visual Question AnsweringAction Recognition