paper-with-me

홈 › Papers

Estimating the Black-box LLM Uncertainty with Distribution-Aligned Adversarial Distillation

2026-05-07 · Huizi Cui, Huan Ma, Qilin Wang, Yuhang Gao, Changqing Zhang arxiv

Large language models (LLMs) have progressed rapidly in complex reasoning and question answering, yet LLM hallucination remains a central bottleneck that hinders practical deployment, especially for commercial black-box LLMs accessible only via APIs. Existing uncertainty quantification methods typically depend on computationally expensive multiple sampling or internal parameters, which prevents real-time estimation and fails to capture information implicit in the black-box reasoning process. To address this issue, we propose Distribution-Aligned Adversarial Distillation (DisAAD), which introduces a generation-discrimination architecture to guide a lightweight proxy model to learn the high-quality regions of the output distribution of the black-box LLM, thus effectively endowing it with the ability to know whether the black-box LLM knows or not. Subsequently, we use the proxy model to reproduce the specific responses of the black-box LLM and estimate the corresponding uncertainty based on evidence learning. Extensive experiments have verified the effectiveness and promise of our proposed method, indicating that a proxy model even one that only accounts for 1\% of the target LLM's size can achieve reliable uncertainty quantification.

📄 PDF Abstract BibTeX arXiv:2605.05777

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Uncertainty-Aware Unsupervised Domain Adaptation in Object Detection

2021-02-27 · Dayan Guan, Jiaxing Huang, Aoran Xiao, Shijian Lu 외

Unsupervised domain adaptive object detection aims to adapt detectors from a labelled source domain to an unlabelled target domain. Most existing works take a two-stage strategy that first generates region proposals and …

Domain AdaptationObjectobject-detectionObject Detection+1

A unifying Bayesian framework for adversarial robustness

2025-10-10 · Pablo G. Arce, Roi Naveiro, David Ríos Insua arxiv

The vulnerability of machine learning models to adversarial attacks remains a critical societal security challenge. Traditional defenses, such as adversarial training, typically robustify models by minimizing a worst-cas…

Adversarial Robustness

Prior Networks for Detection of Adversarial Attacks

2018-12-06 · Andrey Malinin, Mark Gales

Adversarial examples are considered a serious issue for safety critical applications of AI, such as finance, autonomous vehicle control and medicinal applications. Though significant work has resulted in increased robust…

Adversarial AttackAdversarial Attack Detection

Sign Bits Are All You Need for Black-Box Attacks

2020-05-01 · ICLR 2020 1 · Abdullah Al-Dujaili, Una-May O'Reilly

We present a novel black-box adversarial attack algorithm with state-of-the-art model evasion rates for query efficiency under $\ell_\infty$ and $\ell_2$ metrics. It exploits a \textit{sign-based}, rather than magnitude-…

Adversarial AttackAllDimensionality Reduction

Adversarial Robustness on In- and Out-Distribution Improves Explainability

2020-03-20 · ECCV 2020 8 · Maximilian Augustin, Alexander Meinke, Matthias Hein

Neural networks have led to major improvements in image classification but suffer from being non-robust to adversarial changes, unreliable uncertainty estimates on out-distribution samples and their inscrutable black-box…

Adversarial Robustnessimage-classificationImage Classification