paper-with-me

홈 › Papers

LLMProxy: Reducing Cost to Access Large Language Models

2024-10-04 · Noah Martin, Abdullah Bin Faisal, Hiba Eltigani, Rukhshan Haroon, Swaminathan Lamelas, Fahad Dogar

In this paper, we make a case for a proxy for large language models which has explicit support for cost-saving optimizations. We design LLMProxy, which supports three key optimizations: model selection, context management, and caching. These optimizations present tradeoffs in terms of cost, inference time, and response quality, which applications can navigate through our high level, bidirectional interface. As a case study, we implement a WhatsApp-based Q&A service that uses LLMProxy to provide a rich set of features to the users. This service is deployed on a small scale (100+ users) leveraging the cloud; it has been operational for 15+ weeks and users have asked 1400+ questions so far. We report on the experiences of running this service as well as microbenchmark the specific benefits of the various cost-optimizations we present in this paper.

📄 PDF Abstract BibTeX arXiv:2410.11857

Code (0)

등록된 구현이 없습니다.

Tasks

ManagementModel SelectionNavigate

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Advanced Black-Box Tuning of Large Language Models with Limited API Calls

2025-11-13 · Zhikang Xie, Weilin Wan, Peizhu Gong, Weizhong Zhang 외 arxiv

Black-box tuning is an emerging paradigm for adapting large language models (LLMs) to better achieve desired behaviors, particularly when direct access to model parameters is unavailable. Current strategies, however, oft…

Towards Optimizing SQL Generation via LLM Routing

2024-11-06 · Mohammadhossein Malekpour, Nour Shaheen, Foutse khomh, Amine Mhedhbi

Text-to-SQL enables users to interact with databases through natural language, simplifying access to structured data. Although highly capable large language models (LLMs) achieve strong accuracy for complex queries, they…

Text to SQLText-To-SQL

Efficient Agents: Building Effective Agents While Reducing Cost

2025-07-24 · Ningning Wang, Xavier Hu, Pai Liu, He Zhu 외 arxiv

The remarkable capabilities of Large Language Model (LLM)-driven agents have enabled sophisticated systems to tackle complex, multi-step tasks, but their escalating costs threaten scalability and accessibility. This work…

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference

2026-07-22 · Niqi Lyu, Pengtao Shi, Wei Qiu, Jianlin Zhong 외 arxiv

Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a c…

Mathematical Reasoning

A Generative Caching System for Large Language Models

2025-03-22 · Arun Iyengar, Ashish Kundu, Ramana Kompella, Sai Nandan Mamidi

Caching has the potential to be of significant benefit for accessing large language models (LLMs) due to their high latencies which typically range from a small number of seconds to well over a minute. Furthermore, many …