paper-with-me

홈 › Papers

Beyond the Context Window: A Cost-Performance Analysis of Fact-Based Memory vs. Long-Context LLMs for Persistent Agents

2026-03-05 · Natchanon Pollertlam, Witchayut Kornsuwannawit arxiv

Persistent conversational AI systems face a choice between passing full conversation histories to a long-context large language model (LLM) and maintaining a dedicated memory system that extracts and retrieves structured facts. We compare a fact-based memory system built on the Mem0 framework against long-context LLM inference on three memory-centric benchmarks - LongMemEval, LoCoMo, and PersonaMemv2 - and evaluate both architectures on accuracy and cumulative API cost. Long-context GPT-5-mini achieves higher factual recall on LongMemEval and LoCoMo, while the memory system is competitive on PersonaMemv2, where persona consistency depends on stable, factual attributes suited to flat-typed extraction. We construct a cost model that incorporates prompt caching and show that the two architectures have structurally different cost profiles: long-context inference incurs a per-turn charge that grows with context length even under caching, while the memory system's per-turn read cost remains roughly fixed after a one-time write phase. At a context length of 100k tokens, the memory system becomes cheaper after approximately ten interaction turns, with the break-even point decreasing as context length grows. These results characterize the accuracy-cost trade-off between the two approaches and provide a concrete criterion for selecting between them in production deployments.

📄 PDF Abstract BibTeX arXiv:2603.04814

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

2024-02-21 · Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu 외

Large context window is a desirable feature in large language models (LLMs). However, due to high fine-tuning costs, scarcity of long texts, and catastrophic values introduced by new token positions, current extended con…

8k

Exploring Context Window of Large Language Models via Decomposed Positional Vectors

2024-05-28 · Zican Dong, Junyi Li, Xin Men, Wayne Xin Zhao 외

Transformer-based large language models (LLMs) typically have a limited context window, resulting in significant performance degradation when processing text beyond the length of the context window. Extensive studies hav…

MemGPT: Towards LLMs as Operating Systems

2023-10-12 · Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang 외

Large language models (LLMs) have revolutionized AI, but are constrained by limited context windows, hindering their utility in tasks like extended conversations and document analysis. To enable using context beyond limi…

Management

$π$-Attention: Online Efficient Sparse Transformers for Long-Context Modeling

2025-11-12 · Dong Liu, Yanxuan Yu arxiv

Sparse attention is crucial in long-context Transformers, which restricts each token to a limited neighborhood and thereby reduces the quadratic cost of full self-attention. Local windows capture nearby context effective…

Long-range modeling

PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training

2023-09-19 · Dawei Zhu, Nan Yang, Liang Wang, YiFan Song 외

Large Language Models (LLMs) are trained with a pre-defined context length, restricting their use in scenarios requiring long inputs. Previous efforts for adapting LLMs to a longer length usually requires fine-tuning wit…

2kPosition