paper-with-me

Papers

SciTrust 2.0: A Comprehensive Framework for Evaluating Trustworthiness of Large Language Models in Scientific Applications

2025-10-29 · Emily Herron, Junqi Yin, Feiyi Wang arxiv

Large language models (LLMs) have demonstrated transformative potential in scientific research, yet their deployment in high-stakes contexts raises significant trustworthiness concerns. Here, we introduce SciTrust 2.0, a comprehensive framework for evaluating LLM trustworthiness in scientific applications across four dimensions: truthfulness, adversarial robustness, scientific safety, and scientific ethics. Our framework incorporates novel, open-ended truthfulness benchmarks developed through a verified reflection-tuning pipeline and expert validation, alongside a novel ethics benchmark for scientific research contexts covering eight subcategories including dual-use research and bias. We evaluated seven prominent LLMs, including four science-specialized models and three general-purpose industry models, using multiple evaluation metrics including accuracy, semantic similarity measures, and LLM-based scoring. General-purpose industry models overall outperformed science-specialized models across each trustworthiness dimension, with GPT-o4-mini demonstrating superior performance in truthfulness assessments and adversarial robustness. Science-specialized models showed significant deficiencies in logical and ethical reasoning capabilities, along with concerning vulnerabilities in safety evaluations, particularly in high-risk domains such as biosecurity and chemical weapons. By open-sourcing our framework, we provide a foundation for developing more trustworthy AI systems and advancing research on model safety and ethics in scientific contexts.

📄 PDF Abstract BibTeX arXiv:2510.25908

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial RobustnessSemantic Similarity

Similar Papers 제목 키워드 기반

TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health

2026-03-03 · Zixin Xiong, Ziteng Wang, Haotian Fan, Xinjie Zhang 외 arxiv

While Large Language Models (LLMs) demonstrate significant potential in providing accessible mental health support, their practical deployment raises critical trustworthiness concerns due to the domains high-stakes and s…

FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain

2025-10-17 · Tiansheng Hu, Tongyan Hu, Liuyang Bai, Yilun Zhao 외 arxiv

Recent LLMs have demonstrated promising ability in solving finance related problems. However, applying LLMs in real-world finance application remains challenging due to its high risk and high stakes property. This paper …

AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models

2025-05-22 · Kai Li, Can Shen, Yile Liu, Jirui Han 외

The rapid advancement and expanding applications of Audio Large Language Models (ALLMs) demand a rigorous understanding of their trustworthiness. However, systematic research on evaluating these models, particularly conc…

BenchmarkingFairnessHallucination

DuTrust: A Sentiment Analysis Dataset for Trustworthiness Evaluation

2021-08-30 · Lijie Wang, Hao liu, Shuyuan Peng, Hongxuan Tang 외

While deep learning models have greatly improved the performance of most artificial intelligence tasks, they are often criticized to be untrustworthy due to the black-box problem. Consequently, many works have been propo…

Sentiment Analysis

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

2023-08-10 · Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang 외

Ensuring alignment, which refers to making models behave in accordance with human intentions [1,2], has become a critical task before deploying large language models (LLMs) in real-world applications. For instance, OpenA…

FairnessModels Alignment