paper-with-me

홈 › Papers

ODIA: Oriented Distillation for Inline Acceleration of LLM-based Function Calling

2025-07-10 · Hanlong Zhang, Jingsheng Yang, Hao Li, Yuhao He, Franck Gong arxiv

Function Calling is a crucial technique that enables Large Language Models (LLMs) to interact with external systems through APIs. However, the high latency associated with LLM-based Function Calling significantly impacts user experience. This paper presents a novel approach called Oriented Distillation for Inline Acceleration (ODIA) that leverages online user interaction data to accelerate Function Calling. By automatically identifying "simple queries" from production traffic and distilling knowledge from larger models to smaller ones, our method reduces response latency by 45% (expected) and 78% (median) while maintaining accuracy. We demonstrate the effectiveness of our approach through real-world deployment in a music application, where the smaller model successfully handles 60% of traffic with negligible accuracy loss. Our method requires minimal human intervention and continuously improves through automated data collection and model updating, making it a practical solution for production environments.

📄 PDF Abstract BibTeX arXiv:2507.08877

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

IQ-VFI: Implicit Quadratic Motion Estimation for Video Frame Interpolation

2024-01-01 · CVPR 2024 1 · Mengshun Hu, Kui Jiang, Zhihang Zhong, Zheng Wang 외

Advanced video frame interpolation (VFI) algorithms approximate intermediate motions between two input frames to synthesize intermediate frame. However they struggle to handle complex scenarios with curvilinear motio…

Knowledge DistillationMotion EstimationVideo Frame Interpolation

Mainline Automatic Train Horn and Brake Performance Metric

2023-07-05 · Rustam Tagiew

This paper argues for the introduction of a mainline rail-oriented performance metric for driver-replacing on-board perception systems. Perception at the head of a train is divided into several subfunctions. This article…

AUTODIAL: Efficient Asynchronous Task-Oriented Dialogue Model

2023-03-10 · Prajjwal Bhargava, Pooyan Amini, Shahin Shayandeh, Chinnadhurai Sankar

As large dialogue models become commonplace in practice, the problems surrounding high compute requirements for training, inference and larger memory footprint still persists. In this work, we present AUTODIAL, a multi-t…

Dialogue State TrackingmodelPrediction

Hardware Acceleration for Open Radio Access Networks: A Contemporary Overview

2023-05-16 · Lopamudra Kundu, Xingqin Lin, Elena Agostini, Vikrama Ditya

Radio access networks (RAN) are going through a paradigm shift towards interoperable, intelligent, software-defined, and cloud-native open RAN solutions. A key challenge towards the adoption and deployment of open RAN at…

CGoDial: A Large-Scale Benchmark for Chinese Goal-oriented Dialog Evaluation

2022-11-21 · Yinpei Dai, Wanwei He, Bowen Li, Yuchuan Wu 외

Practical dialog systems need to deal with various knowledge sources, noisy user expressions, and the shortage of annotated data. To better solve the above problems, we propose CGoDial, new challenging and comprehensive …

Goal-Oriented DialogRetrieval