paper-with-me

홈 › Papers

Hyperband-based Bayesian Optimization for Black-box Prompt Selection

2024-12-10 · Lennart Schneider, Martin Wistuba, Aaron Klein, Jacek Golebiowski, Giovanni Zappella, Felice Antonio Merra

Optimal prompt selection is crucial for maximizing large language model (LLM) performance on downstream tasks. As the most powerful models are proprietary and can only be invoked via an API, users often manually refine prompts in a black-box setting by adjusting instructions and few-shot examples until they achieve good performance as measured on a validation set. Recent methods addressing static black-box prompt selection face significant limitations: They often fail to leverage the inherent structure of prompts, treating instructions and few-shot exemplars as a single block of text. Moreover, they often lack query-efficiency by evaluating prompts on all validation instances, or risk sub-optimal selection of a prompt by using random subsets of validation instances. We introduce HbBoPs, a novel Hyperband-based Bayesian optimization method for black-box prompt selection addressing these key limitations. Our approach combines a structural-aware deep kernel Gaussian Process to model prompt performance with Hyperband as a multi-fidelity scheduler to select the number of validation instances for prompt evaluations. The structural-aware modeling approach utilizes separate embeddings for instructions and few-shot exemplars, enhancing the surrogate model's ability to capture prompt performance and predict which prompt to evaluate next in a sample-efficient manner. Together with Hyperband as a multi-fidelity scheduler we further enable query-efficiency by adaptively allocating resources across different fidelity levels, keeping the total number of validation instances prompts are evaluated on low. Extensive evaluation across ten benchmarks and three LLMs demonstrate that HbBoPs outperforms state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2412.07820

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian OptimizationLarge Language Model

Methods 이 논문이 사용한 방법론

Gaussian Process Gaussian Processes are non-parametric models for approximating functions. They rely upon a measure of similarity between points (the kernel function) to predict the value for…

Similar Papers 제목 키워드 기반

BOAH: A Tool Suite for Multi-Fidelity Bayesian Optimization & Analysis of Hyperparameters

2019-08-16 · Marius Lindauer, Katharina Eggensperger, Matthias Feurer, André Biedenkapp 외

Hyperparameter optimization and neural architecture search can become prohibitively expensive for regular black-box Bayesian optimization because the training and evaluation of a single model can easily take several hour…

Bayesian OptimizationHyperparameter OptimizationNeural Architecture Search

Combination of Hyperband and Bayesian Optimization for Hyperparameter Optimization in Deep Learning

2018-01-05 · Jiazhuo Wang, Jason Xu, Xuejun Wang

Deep learning has achieved impressive results on many problems. However, it requires high degree of expertise or a lot of experience to tune well the hyperparameters, and such manual tuning process is likely to be biased…

Bayesian OptimizationDeep LearningHyperparameter Optimization

Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization

2016-03-21 · Lisha Li, Kevin Jamieson, Giulia Desalvo, Afshin Rostamizadeh 외

Performance of machine learning algorithms depends critically on identifying a good set of hyperparameters. While recent approaches use Bayesian optimization to adaptively select configurations, we focus on speeding up r…

Bayesian OptimizationHyperparameter Optimization

Bayes Optimal Early Stopping Policies for Black-Box Optimization

2019-02-21 · Matthew Streeter

We derive an optimal policy for adaptively restarting a randomized algorithm, based on observed features of the run-so-far, so as to minimize the expected time required for the algorithm to successfully terminate. Given …

Parameter Optimization with Conscious Allocation (POCA)

2023-12-29 · Joshua Inman, Tanmay Khandait, Giulia Pedrielli, Lalitha Sankar

The performance of modern machine learning algorithms depends upon the selection of a set of hyperparameters. Common examples of hyperparameters are learning rate and the number of layers in a dense neural network. Auto-…