paper-with-me

홈 › Papers

FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance

2025-08-07 · Mengao Zhang, Jiayu Fu, Tanya Warrier, Yuwen Wang, Tianhui Tan, Ke-wei Huang arxiv

Hallucination remains a critical challenge for deploying Large Language Models (LLMs) in finance. Accurate extraction and precise calculation from tabular data are essential for reliable financial analysis, since even minor numerical errors can undermine decision-making and regulatory compliance. Financial applications have unique requirements, often relying on context-dependent, numerical, and proprietary tabular data that existing hallucination benchmarks rarely capture. In this study, we develop a rigorous and scalable framework for evaluating intrinsic hallucinations in financial LLMs, conceptualized as a context-aware masked span prediction task over real-world financial documents. Our main contributions are: (1) a novel, automated dataset creation paradigm using a masking strategy; (2) a new hallucination evaluation dataset derived from S&P 500 annual reports; and (3) a comprehensive evaluation of intrinsic hallucination patterns in state-of-the-art LLMs on financial tabular data. Our work provides a robust methodology for in-house LLM evaluation and serves as a critical step toward building more trustworthy and reliable financial Generative AI systems.

📄 PDF Abstract BibTeX arXiv:2508.05201

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards

2025-05-07 · Manveer Singh Tamber, Forrest Sheng Bao, Chenyu Xu, Ge Luo 외

Hallucinations remain a persistent challenge for LLMs. RAG aims to reduce hallucinations by grounding responses in contexts. However, even when provided context, LLMs still frequently introduce unsupported information or…

BenchmarkingHallucinationHallucination EvaluationRAG

CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models

2025-05-27 · Xiaqiang Tang, Jian Li, Keyu Hu, Du Nan 외

Faithfulness hallucinations are claims generated by a Large Language Model (LLM) not supported by contexts provided to the LLM. Lacking assessment standards, existing benchmarks focus on "factual statements" that rephras…

HallucinationLanguage ModelingLanguage ModellingLarge Language Model

RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation

2024-11-29 · Xianfeng Tan, Yuhan Li, Wenxiang Shang, Yubo Wu 외

Standard clothing asset generation involves creating forward-facing flat-lay garment images displayed on a clear background by extracting clothing information from diverse real-world contexts, which presents significant …

Contrastive LearningRAGRetrievalRetrieval-augmented Generation

Towards Fine-Grained Citation Evaluation in Generated Text: A Comparative Analysis of Faithfulness Metrics

2024-06-21 · Weijia Zhang, Mohammad Aliannejadi, Yifei Yuan, Jiahuan Pei 외

Large language models (LLMs) often produce unsupported or unverifiable content, known as "hallucinations." To mitigate this, retrieval-augmented LLMs incorporate citations, grounding the content in verifiable sources. De…

Binary ClassificationRetrieval

Enhancing Reliability across Short and Long-Form QA via Reinforcement Learning

2025-11-19 · Yudong Wang, Zhe Yang, Wenhan Ma, Zhifang Sui 외 arxiv

While reinforcement learning has unlocked unprecedented complex reasoning in large language models, it has also amplified their propensity for hallucination, creating a critical trade-off between capability and reliabili…

Reinforcement LearningQuestion Answering