paper-with-me

Papers

POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving

2026-04-17 · Shaoang Li, Jian Li arxiv

Edge deployment of large language models (LLMs) increasingly relies on libraries of lightweight LoRA adapters, yet GPU/DRAM can keep only a small resident subset at a time. Serving a request through a non-resident adapter requires paging its weights from storage, incurring measurable latency. This creates a two-timescale online control problem: on a slow timescale, the system selects which adapters remain resident in fast memory, while on a fast timescale it routes each request to an adapter whose context-dependent utility is unknown a priori. The two decisions are tightly coupled: the cache determines the cost of exploration, and the router determines which adapters receive informative feedback. We formulate this joint caching-and-routing problem as a two-timescale contextual bandit and propose POLAR (Paging and Online Learning for Adapter Routing). POLAR pairs a cache-aware LinUCB router with an epoch-based cache controller. We study two variants. A fixed-epoch version provides a robust baseline with worst-case regret guarantees under arbitrary contexts. An epoch-doubling version, POLAR+, adds forced exploration and improved cache optimization to achieve $\widetilde{\mathcal{O}}(d\sqrt{NT}+\sqrt{KT})$ sublinear regret under stochastic regularity and cacheability conditions, where $N$ is the adapter count, $K$ the cache size, $d$ the context dimension, and $T$ the horizon. The routing term matches the standard contextual-bandit rate up to logarithmic factors, showing that the memory hierarchy does not fundamentally slow routing learning. Experiments using 15 real LoRA adapters for Qwen2.5-7B together with measured GPU paging latencies show that adaptive cache control substantially outperforms non-adaptive baselines and exhibits scaling trends consistent with the theory.

📄 PDF Abstract BibTeX arXiv:2604.16583

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems

2025-11-28 · Shashwat Jaiswal, Shrikara Arun, Anjaly Parayil, Ankur Mallick 외 arxiv

Low-Rank Adaptation (LoRA) has become the de facto method for parameter-efficient fine-tuning of large language models (LLMs), enabling rapid adaptation to diverse domains. In production, LoRA-based models are served at …

parameter-efficient fine-tuning

Effective LoRA Adapter Routing using Task Representations

2026-01-29 · Akash Dhasade, Anne-Marie Kermarrec, Igor Pavlovic, Diana Petrescu 외 arxiv

Low-rank adaptation (LoRA) enables parameter efficient specialization of large language models (LLMs) through modular adapters, resulting in rapidly growing public adapter pools spanning diverse tasks. Effectively using …

Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management

2025-04-19 · Hang Zhang, Jiuchen Shi, Yixiao Wang, Quan Chen 외

Multiple Low-Rank Adapters (Multi-LoRAs) are gaining popularity for task-specific Large Language Model (LLM) applications. For multi-LoRA serving, caching hot KV caches and LoRA adapters in high bandwidth memory of accel…

Language ModelingLanguage ModellingLarge Language ModelManagement

MoLoRA: Composable Specialization via Per-Token Adapter Routing

2026-03-16 · Shrey Shah, Justin Wagle arxiv

Multi-adapter serving systems route entire sequences to a single adapter, forcing a choice when requests span multiple domains. This assumption fails in two important settings: (1) multimodal generation, where text and i…

multimodal generation

LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents

2026-02-01 · Hyesung Jeon, Hyeongju Ha, Jae-Joon Kim arxiv

Role specialization in multi-LLM agent systems is often realized via multi-LoRA, where agents share a pretrained backbone and differ only by lightweight adapters. Despite sharing base model weights, each agent independen…