paper-with-me

홈 › Papers

TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models

2025-10-15 · Ruoyu Sun, Da Song, Jiayang Song, Yuheng Huang, Lei Ma arxiv

As Large Language Models (LLMs) continue to revolutionize Natural Language Processing (NLP) applications, critical concerns about their trustworthiness persist, particularly in safety and robustness. To address these challenges, we introduce TRUSTVIS, an automated evaluation framework that provides a comprehensive assessment of LLM trustworthiness. A key feature of our framework is its interactive user interface, designed to offer intuitive visualizations of trustworthiness metrics. By integrating well-known perturbation methods like AutoDAN and employing majority voting across various evaluation methods, TRUSTVIS not only provides reliable results but also makes complex evaluation processes accessible to users. Preliminary case studies on models like Vicuna-7b, Llama2-7b, and GPT-3.5 demonstrate the effectiveness of our framework in identifying safety and robustness vulnerabilities, while the interactive interface allows users to explore results in detail, empowering targeted model improvements. Video Link: https://youtu.be/k1TrBqNVg8g

📄 PDF Abstract BibTeX arXiv:2510.13106

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sampling Preferences Yields Simple Trustworthiness Scores

2025-06-03 · Sean Steinle

With the onset of large language models (LLMs), the performance of artificial intelligence (AI) models is becoming increasingly multi-dimensional. Accordingly, there have been several large, multi-dimensional evaluation …

From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports -- with Preliminary Extension to Lung Cancer

2025-10-27 · Qiuli Wang, Xinhuang Sun, Yonglin Chen, Jie Cheng 외 arxiv

Large language models (LLMs) have demonstrated promising performance in generating diagnostic conclusions from imaging findings, thereby supporting radiology reporting, trainee education, and quality control. However, sy…

Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation

2025-08-21 · Yichi Zhang, Yao Huang, Yifan Wang, Yitong Sun 외 arxiv

The trustworthiness of Multimodal Large Language Models (MLLMs) remains an intense concern despite the significant progress in their capabilities. Existing evaluation and mitigation approaches often focus on narrow aspec…

aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy

2026-08-20 · Fatih Deniz, Yazan Boshmaf, Dorde Popovic, Issa Khalil arxiv

The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or improve across every capability metric …

Classes are not Clusters: Improving Label-based Evaluation of Dimensionality Reduction

2023-08-01 · Hyeon Jeon, Yun-Hsin Kuo, Michaël Aupetit, Kwan-Liu Ma 외

A common way to evaluate the reliability of dimensionality reduction (DR) embeddings is to quantify how well labeled classes form compact, mutually separated clusters in the embeddings. This approach is based on the assu…

Dimensionality Reduction