paper-with-me

Papers

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

2025-01-30 · Spencer Mateega, Carlos Georgescu, Danny Tang

FinanceQA is a testing suite that evaluates LLMs' performance on complex numerical financial analysis tasks that mirror real-world investment work. Despite recent advances, current LLMs fail to meet the strict accuracy requirements of financial institutions, with models failing approximately 60% of realistic tasks that mimic on-the-job analyses at hedge funds, private equity firms, investment banks, and other financial institutions. The primary challenges include hand-spreading metrics, adhering to standard accounting and corporate valuation conventions, and performing analysis under incomplete information - particularly in multi-step tasks requiring assumption generation. This performance gap highlights the disconnect between existing LLM capabilities and the demands of professional financial analysis that are inadequately tested by current testing architectures. Results show that higher-quality training data is needed to support such tasks, which we experiment with using OpenAI's fine-tuning API. FinanceQA is publicly released at this https URL.

📄 PDF Abstract BibTeX arXiv:2501.18062

Code (0)

등록된 구현이 없습니다.

Tasks

Financial Analysis

Similar Papers 제목 키워드 기반

Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning

2025-02-18 · Jingyang Lin, Andy Wong, Tian Xia, Shenghua He 외

Recent advances in Large Language Models (LLMs) have enabled them to process increasingly longer sequences, ranging from 2K to 2M tokens and even beyond. However, simply extending the input sequence length does not neces…

2kLong-Context Understanding

SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities

2025-04-06 · Noga Ben Yoash, Meni Brief, Oded Ovadia, Gil Shenderovitz 외

We introduce SECQUE, a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks. SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categori…

Financial Analysis

Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines

2026-03-09 · Akshay Gulati, Kanha Singhania, Tushar Banga, Parth Arora 외 arxiv

Large language models are increasingly used for financial analysis and investment research, yet systematic evaluation of their financial reasoning capabilities remains limited. In this work, we introduce the AI Financial…

FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models

2024-01-01 · Shu Liu, Shangqing Zhao, Chenghao Jia, Xinlin Zhuang 외

Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of tasks. However, their proficiency and reliability in the specialized domain of financial data analysis, particularly focusing …

Benchmarking

FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis

2025-10-15 · Fengbin Zhu, Xiang Yao Ng, Ziyang Liu, Chang Liu 외 arxiv

Deep Research (DR) agents, powered by advanced Large Language Models (LLMs), have recently garnered increasing attention for their capability in conducting complex research tasks. However, existing literature lacks a rig…