paper-with-me

Papers

BizBench: A Quantitative Reasoning Benchmark for Business and Finance

2023-11-11 · Rik Koncel-Kedziorski, Michael Krumdick, Viet Lai, Varshini Reddy, Charles Lovering, Chris Tanner

Answering questions within business and finance requires reasoning, precision, and a wide-breadth of technical knowledge. Together, these requirements make this domain difficult for large language models (LLMs). We introduce BizBench, a benchmark for evaluating models' ability to reason about realistic financial problems. BizBench comprises eight quantitative reasoning tasks, focusing on question-answering (QA) over financial data via program synthesis. We include three financially-themed code-generation tasks from newly collected and augmented QA data. Additionally, we isolate the reasoning capabilities required for financial QA: reading comprehension of financial text and tables for extracting intermediate values, and understanding financial concepts and formulas needed to calculate complex solutions. Collectively, these tasks evaluate a model's financial background knowledge, ability to parse financial documents, and capacity to solve problems with code. We conduct an in-depth evaluation of open-source and commercial LLMs, comparing and contrasting the behavior of code-focused and language-focused models. We demonstrate that the current bottleneck in performance is due to LLMs' limited business and financial understanding, highlighting the value of a challenging benchmark for quantitative reasoning within this domain.

📄 PDF Abstract BibTeX arXiv:2311.06602

Code (0)

등록된 구현이 없습니다.

Tasks

Code GenerationProgram SynthesisQuestion AnsweringReading Comprehension

Similar Papers 제목 키워드 기반

QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models

2026-01-13 · Zhaolu Kang, Junhao Gong, Wenqing Hu, Shuo Yin 외 arxiv

Large Language Models (LLMs) have shown strong capabilities across many domains, yet their evaluation in financial quantitative tasks remains fragmented and mostly limited to knowledge-centric question answering. We intr…

Reinforcement LearningMathematical ReasoningQuestion Answering

Grants4Companies: Applying Declarative Methods for Recommending and Reasoning About Business Grants in the Austrian Public Administration (System Description)

2024-06-21 · Björn Lellmann, Philipp Marek, Markus Triska

We describe the methods and technologies underlying the application Grants4Companies. The application uses a logic-based expert system to display a list of business grants suitable for the logged-in business. To evaluate…

BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs

2025-05-26 · Guilong Lu, Xuntao Guo, Rongjunchen Zhang, Wenqiao Zhu 외

Large language models excel in general tasks, yet assessing their reliability in logic-heavy, precision-critical domains like finance, law, and healthcare remains challenging. To address this, we introduce BizFinBench, t…

Question Answering

QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs

2025-12-30 · Shupeng Li, Weipeng Lu, Linyun Liu, Chen Lin 외 arxiv

Domain-specific enhancement of Large Language Models (LLMs) within the financial context has long been a focal point of industrial application. While previous models such as BloombergGPT and Baichuan-Finance primarily fo…

SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning

2026-04-21 · Rania Elbadry, Sarfraz Ahmad, Ahmed Heakl, Dani Bouch 외 arxiv

English financial NLP has advanced rapidly through benchmarks targeting earnings analysis, market sentiment, tabular reasoning, and financial question answering, yet Arabic financial NLP remains virtually nonexistent, de…

Sentiment AnalysisQuestion Answering