paper-with-me

홈 › Papers

DisCEdge: Distributed Context Management for Large Language Models at the Edge

2025-11-27 · Mohammadreza Malekabbasi, Minghe Wang, David Bermbach arxiv

Deploying Large Language Model (LLM) services at the edge benefits latency-sensitive and privacy-aware applications. However, the stateless nature of LLMs makes managing user context (e.g., sessions, preferences) across geo-distributed edge nodes challenging. Existing solutions, such as client-side context storage, introduce network latency and bandwidth overhead, undermining edge deployment advantages. We propose DisCEdge, a distributed context management system that stores and replicates user context in tokenized form across edge nodes. By maintaining context as token sequences, our system avoids redundant computation and enables efficient data replication. We evaluate an open-source prototype in a realistic edge environment. DisCEdge improves median response times by up to 14.46% and lowers median inter-node synchronization overhead by up to 15% compared to a raw-text-based system. It also reduces client request sizes by a median of 90% compared to client-side context management, while guaranteeing data consistency.

📄 PDF Abstract BibTeX arXiv:2511.22599

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Data-driven and distributed governance of building facilities management using decentralized autonomous organization, digital twin, and large language models

2026-04-16 · Reachsak Ly, Alireza Shojaei, Xinghua Gao, Philip Agee 외 arxiv

While traditional AI and data-driven facilities management approaches have improved building operational efficiency, they remain constrained by centralized organizational structures that are vulnerable to cyber attacks, …

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference

2025-05-28 · Yue Zhu, Hao Yu, Chen Wang, Zhuoran Liu 외

The increasing adoption of large language models (LLMs) with extended context windows necessitates efficient Key-Value Cache (KVC) management to optimize inference performance. Inference workloads like Retrieval-Augmente…

ManagementRAGRetrieval-augmented Generation

From Specification to Execution: AI Assisted Scientific Workflow Management

2026-06-16 · Komal Thareja, Hamza Safri, Rajiv Mayani, Anirban Mandal 외 arxiv

Scientific workflow management systems (WMS) support scalable and reproducible execution of complex pipelines, but workflow design, implementation, and debugging remain largely manual and require significant expertise. R…

Federated LearningCode Generation

ActiveMem: Distributed Active Memory for Long-Horizon LLM Reasoning

2026-06-09 · Yunhan Jiang, Wenbin Duan, Shasha Guo, Liang Pang 외 arxiv

Memory is essential for enabling large language model (LLM) agents to handle long-horizon reasoning tasks. Existing memory mechanisms are largely centralized, typically organizing retrieved information and interaction hi…

Reinforcement Learning Based Approaches to Adaptive Context Caching in Distributed Context Management Systems

2022-12-22 · Shakthi Weerasinghe, Arkady Zaslavsky, Seng W. Loke, Amin Abken 외

Performance metrics-driven context caching has a profound impact on throughput and response time in distributed context management systems for real-time context queries. This paper proposes a reinforcement learning based…

Managementreinforcement-learningReinforcement Learning (RL)