paper-with-me

Papers

Sampling Preferences Yields Simple Trustworthiness Scores

2025-06-03 · Sean Steinle

With the onset of large language models (LLMs), the performance of artificial intelligence (AI) models is becoming increasingly multi-dimensional. Accordingly, there have been several large, multi-dimensional evaluation frameworks put forward to evaluate LLMs. Though these frameworks are much more realistic than previous attempts which only used a single score like accuracy, multi-dimensional evaluations can complicate decision-making since there is no obvious way to select an optimal model. This work introduces preference sampling, a method to extract a scalar trustworthiness score from multi-dimensional evaluation results by considering the many characteristics of model performance which users value. We show that preference sampling improves upon alternate aggregation methods by using multi-dimensional trustworthiness evaluations of LLMs from TrustLLM and DecodingTrust. We find that preference sampling is consistently reductive, fully reducing the set of candidate models 100% of the time whereas Pareto optimality never reduces the set by more than 50%. Likewise, preference sampling is consistently sensitive to user priors-allowing users to specify the relative weighting and confidence of their preferences-whereas averaging scores is intransigent to the users' prior knowledge.

📄 PDF Abstract BibTeX arXiv:2506.03399

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness

2024-04-29 · Aaron J. Li, Satyapriya Krishna, Himabindu Lakkaraju

The trustworthiness of Large Language Models (LLMs) refers to the extent to which their outputs are reliable, safe, and ethically aligned, and it has become a crucial consideration alongside their cognitive performance. …

EthicsLanguage Modelling

Enhancing Mutual Trustworthiness in Federated Learning for Data-Rich Smart Cities

2024-05-01 · Osama Wehbi, Sarhad Arisdakessian, Mohsen Guizani, Omar Abdel Wahab 외

Federated learning is a promising collaborative and privacy-preserving machine learning approach in data-rich smart cities. Nevertheless, the inherent heterogeneity of these urban environments presents a significant chal…

Federated LearningPrivacy Preserving

On Verbalized Confidence Scores for LLMs

2024-12-19 · Daniel Yang, Yao-Hung Hubert Tsai, Makoto Yamada

The rise of large language models (LLMs) and their tight integration into our daily life make it essential to dedicate efforts towards their trustworthiness. Uncertainty quantification for LLMs can establish more human t…

Uncertainty Quantification

Trust, but Verify: Using Self-Supervised Probing to Improve Trustworthiness

2023-02-06 · Ailin Deng, Shen Li, Miao Xiong, Zhirui Chen 외

Trustworthy machine learning is of primary importance to the practical deployment of deep learning models. While state-of-the-art models achieve astonishingly good performance in terms of accuracy, recent literature reve…

Out-of-Distribution Detection

A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance

2026-07-17 · Andrea Ferrario arxiv

AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, whether observed changes are tolerable, and how such judgments should be documented in a transparent and …