paper-with-me

홈 › Papers

From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs

2025-04-18 · Jiliang Ni, Jiachen Pu, Zhongyi Yang, Kun Zhou, Hui Wang, Xiaoliang Xiao, Dakui Wang, Xin Li, Jingfeng Luo, Conggang Hu

Large Language Models (LLMs) have significantly advanced artificial intelligence by optimizing traditional Natural Language Processing (NLP) workflows, facilitating their integration into various systems. Many such NLP systems, including ours, directly incorporate LLMs. However, this approach either results in expensive costs or yields suboptimal performance after fine-tuning. In this paper, we introduce a three-stage cost-efficient end-to-end LLM deployment pipeline, comprising prototyping, knowledge transfer, and model compression, to effectively tackle the cost-performance dilemma in LLM-based frameworks. Its high cost-efficiency is manifested not only in simplifying system complexity and producing super-tiny online models with enhanced performance and reduced costs in the results, but also in addressing development cycle constraints, the lack of extensive high-quality data, and limited computational resources during the project development process. In the first stage, we construct an optimal performance prototype system by transforming complex tasks into a function call-based LLM-driven pipeline, which serves as a teacher model to generate high-quality data. In the second stage, we combine techniques like rejection sampling fine-tuning, reinforcement learning, and knowledge distillation to transfer knowledge to 0.5B student models, delivering effective performance at minimal cost. In the final stage, we further compress models to 0.4B via quantization and pruning, achieving ultra-low latency and cost. Extensive experimental results and the framework's modular design suggest cross-domain capabilities and potential applicability in other NLP areas.

📄 PDF Abstract BibTeX arXiv:2504.13471

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationModel CompressionQuantizationTransfer Learning

Methods 이 논문이 사용한 방법론

Pruning 설명 없음
Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Rethinking Optimization and Architecture for Tiny Language Models

2024-02-05 · Yehui Tang, Kai Han, Fangcheng Liu, Yunsheng Ni 외

The power of large language models (LLMs) has been demonstrated through numerous data and computing resources. However, the application of language models on mobile devices is facing huge challenge on the computation and…

Language Modelling

Cross-model Control: Improving Multiple Large Language Models in One-time Training

2024-10-23 · Jiayi Wu, Hao Sun, Hengyi Cai, Lixin Su 외

The number of large language models (LLMs) with varying parameter scales and vocabularies is increasing. While they deliver powerful performance, they also face a set of common optimization needs to meet specific require…

Instruction FollowingLanguage ModelingLanguage Modelling

Can LLMs Revolutionize the Design of Explainable and Efficient TinyML Models?

2025-04-13 · Christophe El Zeinaty, Wassim Hamidouche, Glenn Herrou, Daniel Menard 외

This paper introduces a novel framework for designing efficient neural network architectures specifically tailored to tiny machine learning (TinyML) platforms. By leveraging large language models (LLMs) for neural archit…

Computational EfficiencyEfficient Neural NetworkKnowledge DistillationNeural Architecture Search

STAR: Similarity-guided Teacher-Assisted Refinement for Super-Tiny Function Calling Models

2026-02-03 · Jiliang Ni, Jiachen Pu, Zhongyi Yang, Jingfeng Luo 외 arxiv

The proliferation of Large Language Models (LLMs) in function calling is pivotal for creating advanced AI agents, yet their large scale hinders widespread adoption, necessitating transferring their capabilities into smal…

Knowledge Distillation

TRIM: Token Reduction and Inference Modeling for Cost-Effective Language Generation

2024-12-10 · Alfredo Garrachón Ruiz, Tomás de la Rosa, Daniel Borrajo

The inference cost of Large Language Models (LLMs) is a significant challenge due to their computational demands, specially on tasks requiring long outputs. However, natural language often contains redundancy, which pres…

General KnowledgeText GenerationToken Reduction