paper-with-me

Papers

A Context-Aware Dual-Metric Framework for Confidence Estimation in Large Language Models

2025-08-01 · Mingruo Yuan, Shuyi Zhang, Ben Kao arxiv

Accurate confidence estimation is essential for trustworthy large language models (LLMs) systems, as it empowers the user to determine when to trust outputs and enables reliable deployment in safety-critical applications. Current confidence estimation methods for LLMs neglect the relevance between responses and contextual information, a crucial factor in output quality evaluation, particularly in scenarios where background knowledge is provided. To bridge this gap, we propose CRUX (Context-aware entropy Reduction and Unified consistency eXamination), the first framework that integrates context faithfulness and consistency for confidence estimation via two novel metrics. First, contextual entropy reduction represents data uncertainty with the information gain through contrastive sampling with and without context. Second, unified consistency examination captures potential model uncertainty through the global consistency of the generated answers with and without context. Experiments across three benchmark datasets (CoQA, SQuAD, QuAC) and two domain-specific datasets (BioASQ, EduQG) demonstrate CRUX's effectiveness, achieving the highest AUROC than existing baselines.

📄 PDF Abstract BibTeX arXiv:2508.00600

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CoCoA: Confidence and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models

2025-08-25 · Anant Khandelwal, Manish Gupta, Puneet Agrawal arxiv

Faithful generation in large language models (LLMs) is challenged by knowledge conflicts between parametric memory and external context. Existing contrastive decoding methods tuned specifically to handle conflict often l…

Question Answering

Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning

2025-12-14 · Haiyang Zheng, Nan Pu, Wenjing Li, Teng Long 외 arxiv

The proliferation of synthetic facial imagery has intensified the need for robust Open-World DeepFake Attribution (OW-DFA), which aims to attribute both known and unknown forgeries using labeled data for known types and …

Uncertainty-Aware Post-Hoc Calibration: Mitigating Confidently Incorrect Predictions Beyond Calibration Metrics

2025-10-19 · Hassan Gharoun, Mohammad Sadegh Khorshidi, Kasra Ranjbarigderi, Fang Chen 외 arxiv

Despite extensive research on neural network calibration, existing methods typically apply global transformations that treat all predictions uniformly, overlooking the heterogeneous reliability of individual predictions.…

Semantic Similarity

RAD: Retrieval-Augmented Monocular Metric Depth Estimation for Underrepresented Classes

2026-02-10 · Michael Baltaxe, Dan Levi, Sagie Benaim arxiv

Monocular Metric Depth Estimation (MMDE) is essential for physically intelligent systems, yet accurate depth estimation for underrepresented classes in complex scenes remains a persistent challenge. To address this, we p…

Depth Estimation

GeoCFNet: Geometry-Aware Confidence Field Network for Robot-Assisted Endoscopic Submucosal Dissection

2026-06-11 · Rui Tang, Guankun Wang, Long Bai, Haochen Yin 외 arxiv

Advanced surgical robotics has made robot-assisted endoscopic submucosal dissection (ESD) a promising approach for the en-bloc resection of large lesions, with the potential to reduce recurrence and improve long-term out…