paper-with-me

홈 › Papers

Evaluating Large Language Models for Financial Reasoning: A CFA-Based Benchmark Study

2025-08-29 · Xuan Yao, Qianteng Wang, Xinbo Liu, Ke-Wei Huang arxiv

The rapid advancement of large language models presents significant opportunities for financial applications, yet systematic evaluation in specialized financial contexts remains limited. This study presents the first comprehensive evaluation of state-of-the-art LLMs using 1,560 multiple-choice questions from official mock exams across Levels I-III of CFA, most rigorous professional certifications globally that mirror real-world financial analysis complexity. We compare models distinguished by core design priorities: multi-modal and computationally powerful, reasoning-specialized and highly accurate, and lightweight efficiency-optimized. We assess models under zero-shot prompting and through a novel Retrieval-Augmented Generation pipeline that integrates official CFA curriculum content. The RAG system achieves precise domain-specific knowledge retrieval through hierarchical knowledge organization and structured query generation, significantly enhancing reasoning accuracy in professional financial certification evaluation. Results reveal that reasoning-oriented models consistently outperform others in zero-shot settings, while the RAG pipeline provides substantial improvements particularly for complex scenarios. Comprehensive error analysis identifies knowledge gaps as the primary failure mode, with minimal impact from text readability. These findings provide actionable insights for LLM deployment in finance, offering practitioners evidence-based guidance for model selection and cost-performance optimization.

📄 PDF Abstract BibTeX arXiv:2509.04468

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BizBench: A Quantitative Reasoning Benchmark for Business and Finance

2023-11-11 · Rik Koncel-Kedziorski, Michael Krumdick, Viet Lai, Varshini Reddy 외

Answering questions within business and finance requires reasoning, precision, and a wide-breadth of technical knowledge. Together, these requirements make this domain difficult for large language models (LLMs). We intro…

Code GenerationProgram SynthesisQuestion AnsweringReading Comprehension

FIND: Toward Multimodal Financial Reasoning and Question Answering for Indic Languages

2026-05-13 · Sarmistha Das, Vaibhav Vishal, Syed Ibrahim Ahmad, Manish Gupta 외 arxiv

Financial decision-making in multilingual settings demands accurate numerical reasoning grounded in diverse modalities, yet existing benchmarks largely overlook this high-stakes, real-world challenge, especially for Indi…

Multimodal ReasoningQuestion Answering

Advanced Financial Reasoning at Scale: A Comprehensive Evaluation of Large Language Models on CFA Level III

2025-06-29 · Pranam Shetty, Abhisek Upadhayaya, Parth Mitesh Shah, Srikanth Jagabathula 외

As financial institutions increasingly adopt Large Language Models (LLMs), rigorous domain-specific evaluation becomes critical for responsible deployment. This paper presents a comprehensive benchmark evaluating 23 stat…

Model SelectionMultiple-choice

FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation

2025-12-31 · Zichen Tang, Haihong E, Rongjin Li, Jiacheng Liu 외 arxiv

We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to existing benchmarks, our work delivers three…

Multimodal Reasoning

Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines

2026-03-09 · Akshay Gulati, Kanha Singhania, Tushar Banga, Parth Arora 외 arxiv

Large language models are increasingly used for financial analysis and investment research, yet systematic evaluation of their financial reasoning capabilities remains limited. In this work, we introduce the AI Financial…