paper-with-me

홈 › Papers

Evaluating Large Language Models (LLMs) in Financial NLP: A Comparative Study on Financial Report Analysis

2025-07-24 · Md Talha Mohsin arxiv

Large language models (LLMs) are increasingly used to support the analysis of complex financial disclosures, yet their reliability, behavioral consistency, and transparency remain insufficiently understood in high-stakes settings. This paper presents a controlled evaluation of five transformer-based LLMs applied to question answering over the Business sections of U.S. 10-K filings. To capture complementary aspects of model behavior, we combine human evaluation, automated similarity metrics, and behavioral diagnostics under standardized and context-controlled prompting conditions. Human assessments indicate that models differ in their average performance across qualitative dimensions such as relevance, completeness, clarity, conciseness, and factual accuracy, though inter-rater agreement is modest, reflecting the subjective nature of these criteria. Automated metrics reveal systematic differences in lexical overlap and semantic similarity across models, while behavioral diagnostics highlight variation in response stability and cross-prompt alignment. Importantly, no single model consistently dominates across all evaluation perspectives. Together, these findings suggest that apparent performance differences should be interpreted as relative tendencies under the tested conditions rather than definitive indicators of general reliability. The results underscore the need for evaluation frameworks that account for human disagreement, behavioral variability, and interpretability when deploying LLMs in financially consequential applications.

📄 PDF Abstract BibTeX arXiv:2507.22936

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SimilarityQuestion Answering

Similar Papers 제목 키워드 기반

Will LLMs be Professional at Fund Investment? DeepFund: A Live Arena Perspective

2025-03-24 · Changlun Li, Yao Shi, Yuyu Luo, Nan Tang

Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, but their effectiveness in financial decision-making remains inadequately evaluated. Current benchmarks primarily assess LLMs…

Decision Making

Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models

2024-11-09 · XiaoJun Wu, Junxi Liu, Huanyi Su, Zhouchi Lin 외

As large language models become increasingly prevalent in the financial sector, there is a pressing need for a standardized method to comprehensively assess their performance. However, existing finance benchmarks often s…

Evaluating Large Language Models on Financial Report Summarization: An Empirical Study

2024-11-11 · Xinqi Yang, Scott Zang, Yong Ren, Dingjie Peng 외

In recent years, Large Language Models (LLMs) have demonstrated remarkable versatility across various applications, including natural language understanding, domain-specific knowledge tasks, etc. However, applying LLMs t…

Natural Language Understanding

Comparing LLMs for Sentiment Analysis in Financial Market News

2025-10-03 · Lucas Eduardo Pereira Teles, Carlos M. S. Figueiredo arxiv

This article presents a comparative study of large language models (LLMs) in the task of sentiment analysis of financial market news. This work aims to analyze the performance difference of these models in this important…

Sentiment Analysis

Is ChatGPT a Financial Expert? Evaluating Language Models on Financial Natural Language Processing

2023-10-19 · Yue Guo, Zian Xu, Yi Yang

The emergence of Large Language Models (LLMs), such as ChatGPT, has revolutionized general natural language preprocessing (NLP) tasks. However, their expertise in the financial domain lacks a comprehensive evaluation. To…

DecoderLanguage Model EvaluationLanguage ModelingLanguage Modelling