Teola: Towards End-to-End Optimization of LLM-based Applications
Large language model (LLM)-based applications consist of both LLM and non-LLM components, each contributing to the end-to-end latency. Despite great efforts to optimize LLM inference, end-to-end workflow optimization has been overlooked. Existing frameworks employ coarse-grained orchestration with task modules, which confines optimizations to within each module and yields suboptimal scheduling decisions. We propose fine-grained end-to-end orchestration, which utilizes task primitives as the basic units and represents each query's workflow as a primitive-level dataflow graph. This explicitly exposes a much larger design space, enables optimizations in parallelization and pipelining across primitives of different modules, and enhances scheduling to improve application-level performance. We build Teola, a novel orchestration framework for LLM-based applications that implements this scheme. Comprehensive experiments show that Teola can achieve up to 2.09x speedup over existing systems across various popular LLM applications. The code is available at https://github.com/NetX-lab/Ayo.
Code (1)
Tasks
Language ModelingLanguage ModellingLarge Language ModelSchedulingSimilar Papers 제목 키워드 기반
Multi-objective optimization of energy consumption and execution time in a single level cache memory for embedded systems
Current embedded systems are specifically designed to run multimedia applications. These applications have a big impact on both performance and energy consumption. Both metrics can be optimized selecting the best cache c…
Evolutionary AlgorithmsGraph Bayesian Optimization: Algorithms, Evaluations and Applications
Network structure optimization is a fundamental task in complex network analysis. However, almost all the research on Bayesian optimization is aimed at optimizing the objective functions with vectorial inputs. In this wo…
Bayesian OptimizationConvex Quaternion Optimization for Signal Processing: Theory and Applications
Convex optimization methods have been extensively used in the fields of communications and signal processing. However, the theory of quaternion optimization is currently not as fully developed and systematic as that of c…
Simplifying deflation for non-convex optimization with applications in Bayesian inference and topology optimization
Non-convex optimization problems have multiple local optimal solutions. Non-convex optimization problems are commonly found in numerous applications. One of the methods recently proposed to efficiently explore multiple l…
Bayesian InferenceLow-Complexity Particle Swarm Optimization for Time-Critical Applications
Particle swam optimization (PSO) is a popular stochastic optimization method that has found wide applications in diverse fields. However, PSO suffers from high computational complexity and slow convergence speed. High co…
Stochastic Optimization