paper-with-me

Papers

VELO: A Vector Database-Assisted Cloud-Edge Collaborative LLM QoS Optimization Framework

2024-06-19 · Zhi Yao, Zhiqing Tang, Jiong Lou, Ping Shen, Weijia Jia

The Large Language Model (LLM) has gained significant popularity and is extensively utilized across various domains. Most LLM deployments occur within cloud data centers, where they encounter substantial response delays and incur high costs, thereby impacting the Quality of Services (QoS) at the network edge. Leveraging vector database caching to store LLM request results at the edge can substantially mitigate response delays and cost associated with similar requests, which has been overlooked by previous research. Addressing these gaps, this paper introduces a novel Vector database-assisted cloud-Edge collaborative LLM QoS Optimization (VELO) framework. Firstly, we propose the VELO framework, which ingeniously employs vector database to cache the results of some LLM requests at the edge to reduce the response time of subsequent similar requests. Diverging from direct optimization of the LLM, our VELO framework does not necessitate altering the internal structure of LLM and is broadly applicable to diverse LLMs. Subsequently, building upon the VELO framework, we formulate the QoS optimization problem as a Markov Decision Process (MDP) and devise an algorithm grounded in Multi-Agent Reinforcement Learning (MARL) to decide whether to request the LLM in the cloud or directly return the results from the vector database at the edge. Moreover, to enhance request feature extraction and expedite training, we refine the policy network of MARL and integrate expert demonstrations. Finally, we implement the proposed algorithm within a real edge system. Experimental findings confirm that our VELO framework substantially enhances user satisfaction by concurrently diminishing delay and resource consumption for edge users utilizing LLMs.

📄 PDF Abstract BibTeX arXiv:2406.13399

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingLarge Language ModelMulti-agent Reinforcement Learning

Similar Papers 제목 키워드 기반

A 3D Motion Vector Database for Dynamic Point Clouds

2020-08-19

Due to the large amount of data that point clouds represent and the differences in geometry of successive frames, the generation of motion vectors for an entire point cloud dataset may require a significant amount of tim…

Motion Estimation

Assisted RTF-Vector-Based Binaural Direction of Arrival Estimation Exploiting a Calibrated External Microphone Array

2022-11-30 · Daniel Fejgin, Simon Doclo

Recently, a relative transfer function (RTF)-vector-based method has been proposed to estimate the direction of arrival (DOA) of a target speaker for a binaural hearing aid setup, assuming the availability of external mi…

Direction of Arrival Estimation

CABLE: Cloud-Assisted Bandwidth-efficient LMM-based Encoding for V2X Systems

2026-06-17 · Haohua Que, Zhipeng Bao, Qianyi Wu, Handong Yao arxiv

Cloud-hosted large multimodal models (LMMs) can provide strong open-vocabulary perception for Vehicle-to-Everything systems, but naively transmitting full-resolution frames from edge to cloud causes severe communication …

Bang for the Buck: Vector Search on Cloud CPUs

2025-05-12 · Leonardo Kuffo, Peter Boncz

Vector databases have emerged as a new type of systems that support efficient querying of high-dimensional vectors. Many of these offer their database as a service in the cloud. However, the variety of available CPUs and…

CPUQuantization

Cost-Effective, Low Latency Vector Search with Azure Cosmos DB

2025-05-09 · Nitish Upreti, Krishnan Sundaram, Hari Sudan Sundar, Samer Boshra 외

Vector indexing enables semantic search over diverse corpora and has become an important interface to databases for both users and AI agents. Efficient vector search requires deep optimizations in database systems. This …