paper-with-me

홈 › Papers

MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning

2026-03-03 · Jiejun Tan, Zhicheng Dou, Liancheng Zhang, Yuyang Hu, Yiruo Cheng, Ji-Rong Wen arxiv

As Large Language Models (LLMs) are increasingly used for long-duration tasks, maintaining effective long-term memory has become a critical challenge. Current methods often face a trade-off between cost and accuracy. Simple storage methods often fail to retrieve relevant information, while complex indexing methods (such as memory graphs) require heavy computation and can cause information loss. Furthermore, relying on the working LLM to process all memories is computationally expensive and slow. To address these limitations, we propose MemSifter, a novel framework that offloads the memory retrieval process to a small-scale proxy model. Instead of increasing the burden on the primary working LLM, MemSifter uses a smaller model to reason about the task before retrieving the necessary information. This approach requires no heavy computation during the indexing phase and adds minimal overhead during inference. To optimize the proxy model, we introduce a memory-specific Reinforcement Learning (RL) training paradigm. We design a task-outcome-oriented reward based on the working LLM's actual performance in completing the task. The reward measures the actual contribution of retrieved memories by mutiple interactions with the working LLM, and discriminates retrieved rankings by stepped decreasing contributions. Additionally, we employ training techniques such as Curriculum Learning and Model Merging to improve performance. We evaluated MemSifter on eight LLM memory benchmarks, including Deep Research tasks. The results demonstrate that our method meets or exceeds the performance of existing state-of-the-art approaches in both retrieval accuracy and final task completion. MemSifter offers an efficient and scalable solution for long-term LLM memory. We have open-sourced the model weights, code, and training data to support further research.

📄 PDF Abstract BibTeX arXiv:2603.03379

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference

2026-01-30 · Nikhil Gopal, Kostis Kaffes arxiv

Large Language Model (LLM) inference is increasingly constrained by GPU memory capacity rather than compute throughput, driven by growing model sizes and the linear growth of the key-value (KV) cache during autoregressiv…

MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall

2025-09-02 · Avinash Maurya, M. Mustafa Rafique, Franck Cappello, Bogdan Nicolae arxiv

Training LLMs larger than the aggregated memory of multiple GPUs is increasingly necessary due to the faster growth of LLM sizes compared to GPU memory. To this end, multi-tier host memory or disk offloading techniques a…

Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage

2025-06-06 · Ziqi Yuan, Haoyang Zhang, Yirui Eric Zhou, Apoorve Mohan 외

We present the design and implementation of a new lifetime-aware tensor offloading framework for GPU memory expansion using low-cost PCIe-based solid-state drives (SSDs). Our framework, TERAIO, is developed explicitly fo…

CPUGPULarge Language Model

Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors

2024-06-14 · Siyuan Chen, Zhuofeng Wang, Zelong Guan, Yudong Liu 외

Fine-tuning large language models (LLMs) requires significant memory, often exceeding the capacity of a single GPU. A common solution to this memory challenge is offloading compute and data from the GPU to the CPU. Howev…

CPUGPU

Retrieval-Augmented Generation for Mobile Edge Computing via Large Language Model

2024-12-30 · Runtao Ren, Yinyu Wu, Xuhui Zhang, Jinke Ren 외

The rapid evolution of mobile edge computing (MEC) has introduced significant challenges in optimizing resource allocation in highly dynamic wireless communication systems, in which task offloading decisions should be ma…

Edge-computingInformation RetrievalLanguage ModelingLanguage Modelling+4