Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
Large Language Model (LLM) services exhibit impressive capability on unlearned tasks leveraging only a few examples by in-context learning (ICL). However, the success of ICL varies depending on the task and context, leading to heterogeneous service quality. Directly estimating the performance of LLM services at each invocation can be laborious, especially requiring abundant labeled data or internal information within the LLM. This paper introduces a novel method to estimate the performance of LLM services across different tasks and contexts, which can be "plug-and-play" utilizing only a few unlabeled samples like ICL. Our findings suggest that the negative log-likelihood and perplexity derived from LLM service invocation can function as effective and significant features. Based on these features, we utilize four distinct meta-models to estimate the performance of LLM services. Our proposed method is compared against unlabeled estimation baselines across multiple LLM services and tasks. And it is experimentally applied to two scenarios, demonstrating its effectiveness in the selection and further optimization of LLM services.
Code (1)
Tasks
In-Context LearningLanguage ModelingLanguage ModellingLarge Language ModelMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Secure and Scalable Network Slicing with Plug-and-Play Support for Power Distribution System Communication Networks
With the rapid development of power distribution systems (PDSs), the number of terminal devices and the types of delivered services involved are constantly growing. These trends make the operations of PDSs highly depende…
CurvPnP: Plug-and-play Blind Image Restoration with Deep Curvature Denoiser
Due to the development of deep learning-based denoisers, the plug-and-play strategy has achieved great success in image restoration problems. However, existing plug-and-play image restoration methods are designed for non…
DeblurringDenoisingImage DenoisingImage Restoration+3Vehicle-to-grid plug-in forecasting for participation in ancillary services markets
Electric vehicle (EV) charge points (CPs) can be used by aggregators to provide frequency response (FR) services. Aggregators must have day-ahead half-hourly forecasts of minimum aggregate vehicle-to-grid (V2G) plug-in t…
SING: A Plug-and-Play DNN Learning Technique
We propose SING (StabIlized and Normalized Gradient), a plug-and-play technique that improves the stability and generalization of the Adam(W) optimizer. SING is straightforward to implement and has minimal computational …
Depth Estimationimage-classificationImage ClassificationDirected Beam Search: Plug-and-Play Lexically Constrained Language Generation
Large pre-trained language models are capable of generating realistic text. However, controlling these models so that the generated text satisfies lexical constraints, i.e., contains specific words, is a challenging prob…
Language ModelingLanguage ModellingMachine TranslationStory Generation+2