paper-with-me

Papers

FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain

2025-10-17 · Tiansheng Hu, Tongyan Hu, Liuyang Bai, Yilun Zhao, Arman Cohan, Chen Zhao arxiv

Recent LLMs have demonstrated promising ability in solving finance related problems. However, applying LLMs in real-world finance application remains challenging due to its high risk and high stakes property. This paper introduces FinTrust, a comprehensive benchmark specifically designed for evaluating the trustworthiness of LLMs in finance applications. Our benchmark focuses on a wide range of alignment issues based on practical context and features fine-grained tasks for each dimension of trustworthiness evaluation. We assess eleven LLMs on FinTrust and find that proprietary models like o4-mini outperforms in most tasks such as safety while open-source models like DeepSeek-V3 have advantage in specific areas like industry-level fairness. For challenging task like fiduciary alignment and disclosure, all LLMs fall short, showing a significant gap in legal awareness. We believe that FinTrust can be a valuable benchmark for LLMs' trustworthiness evaluation in finance domain.

📄 PDF Abstract BibTeX arXiv:2510.15232

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

2023-06-20 · NeurIPS 2023 11 · Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie 외

Generative Pre-trained Transformer (GPT) models have exhibited exciting progress in their capabilities, capturing the interest of practitioners and the public alike. Yet, while the literature on the trustworthiness of GP…

Adversarial RobustnessEthicsFairness

Measuring Consistency in Text-based Financial Forecasting Models

2023-05-15 · Linyi Yang, Yingpeng Ma, Yue Zhang

Financial forecasting has been an important and active area of machine learning research, as even the most modest advantage in predictive accuracy can be parlayed into significant financial gains. Recent advances in natu…

XTRUST: On the Multilingual Trustworthiness of Large Language Models

2024-09-24 · Yahan Li, Yi Wang, Yi Chang, Yuan Wu

Large language models (LLMs) have demonstrated remarkable capabilities across a range of natural language processing (NLP) tasks, capturing the attention of both practitioners and the broader public. A key question that …

EthicsFairnessHallucinationMisinformation

AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic

2024-03-14 · Emad A. Alghamdi, Reem I. Masoud, Deema Alnuhait, Afnan Y. Alomairi 외

The swift progress and widespread acceptance of artificial intelligence (AI) systems highlight a pressing requirement to comprehend both the capabilities and potential risks associated with AI. Given the linguistic compl…

EthicsMultiple-choice

Agentar-Fin-R1: Enhancing Financial Intelligence through Domain Expertise, Training Efficiency, and Advanced Reasoning

2025-07-22 · Yanjun Zheng, Xiyang Du, Longfei Liao, Xiaoke Zhao 외 arxiv

Large Language Models (LLMs) exhibit considerable promise in financial applications; however, prevailing models frequently demonstrate limitations when confronted with scenarios that necessitate sophisticated reasoning c…