paper-with-me

홈 › Papers

Benchmarking Large Language Model Uncertainty for Prompt Optimization

2024-09-16 · Pei-Fu Guo, Yun-Da Tsai, Shou-De Lin

Prompt optimization algorithms for Large Language Models (LLMs) excel in multi-step reasoning but still lack effective uncertainty estimation. This paper introduces a benchmark dataset to evaluate uncertainty metrics, focusing on Answer, Correctness, Aleatoric, and Epistemic Uncertainty. Through analysis of models like GPT-3.5-Turbo and Meta-Llama-3.1-8B-Instruct, we show that current metrics align more with Answer Uncertainty, which reflects output confidence and diversity, rather than Correctness Uncertainty, highlighting the need for improved metrics that are optimization-objective-aware to better guide prompt optimization. Our code and dataset are available at https://github.com/0Frett/PO-Uncertainty-Benchmarking.

📄 PDF Abstract BibTeX arXiv:2409.10044

Code (1)

0frett/po-uncertainty-benchmarking 공식 구현

Tasks

BenchmarkingDiversityLanguage ModelingLanguage ModellingLarge Language Modelmodel

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs

2025-05-30 · Gabrielle Kaili-May Liu, Gal Yona, Avi Caciularu, Idan Szpektor 외

A critical component in the trustworthiness of LLMs is reliable uncertainty communication, yet LLMs often use assertive language when conveying false claims, leading to over-reliance and eroded trust. We present the firs…

Benchmarking

How Confident Is the First Token? An Uncertainty-Calibrated Prompt Optimization Framework for Large Language Model Classification and Understanding

2026-02-23 · Wei Chen, Guoyang Ju, Yuanyuan Qi arxiv

With the widespread adoption of large language models (LLMs) in natural language processing, prompt engineering and retrieval-augmented generation (RAG) have become mainstream to enhance LLMs' performance on complex task…

Prompt Engineering

On the Stability of Prompt Ranking in Large Language Model Evaluation

2026-06-23 · Shaoshuai Du, Penghao Liang, Yixian Shen, Chuanqi Shi 외 arxiv

Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evaluated and the top-ranked one is selected for downstream use. This workflow implicit…

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

2025-05-18 · Minghan Chen, Guikun Chen, Wenguan Wang, Yi Yang

Large language models (LLMs) exhibit varying levels of confidence across input prompts (questions): some lead to consistent, semantically similar answers, while others yield diverse or contradictory outputs. This variati…

MathMathematical Reasoning

promptolution: A Unified, Modular Framework for Prompt Optimization

2025-12-02 · Tom Zehle, Timo Heiß, Moritz Schlager, Matthias Aßenmacher 외 arxiv

Prompt optimization has become crucial for enhancing the performance of large language models (LLMs) across a broad range of tasks. Although many research papers demonstrate its effectiveness, practical adoption is hinde…