paper-with-me

홈 › Papers

RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?

2026-02-06 · Yuyang Dai, Yan Lin, Zhuohan Xie, Yuxia Wang arxiv

Reliable financial reasoning requires knowing not only how to answer, but also when an answer cannot be justified. In real financial practice, problems often rely on implicit assumptions that are taken for granted rather than stated explicitly, causing problems to appear solvable while lacking enough information for a definite answer. We introduce REALFIN, a bilingual benchmark that evaluates financial reasoning by systematically removing essential premises from exam-style questions while keeping them linguistically plausible. Based on this, we evaluate models under three formulations that test answering, recognizing missing information, and rejecting unjustified options, and find consistent performance drops when key conditions are absent. General-purpose models tend to over-commit and guess, while most finance-specialized models fail to clearly identify missing premises. These results highlight a critical gap in current evaluations and show that reliable financial models must know when a question should not be answered.

📄 PDF Abstract BibTeX arXiv:2602.07096

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BizBench: A Quantitative Reasoning Benchmark for Business and Finance

2023-11-11 · Rik Koncel-Kedziorski, Michael Krumdick, Viet Lai, Varshini Reddy 외

Answering questions within business and finance requires reasoning, precision, and a wide-breadth of technical knowledge. Together, these requirements make this domain difficult for large language models (LLMs). We intro…

Code GenerationProgram SynthesisQuestion AnsweringReading Comprehension

FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains

2023-11-16 · Yilun Zhao, Hongjun Liu, Yitao Long, Rui Zhang 외

We introduce FinanceMath, a novel benchmark designed to evaluate LLMs' capabilities in solving knowledge-intensive math reasoning problems. Compared to prior works, this study features three core advancements. First, Fin…

MathMath Word Problem SolvingRetrieval

MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning

2024-11-05 · Ziliang Gan, Yu Lu, Dong Zhang, Haohan Li 외

In recent years, multimodal benchmarks for general domains have guided the rapid development of multimodal models on general tasks. However, the financial field has its peculiarities. It features unique graphical images …

MMEQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning

2025-02-18 · Jingyang Lin, Andy Wong, Tian Xia, Shenghua He 외

Recent advances in Large Language Models (LLMs) have enabled them to process increasingly longer sequences, ranging from 2K to 2M tokens and even beyond. However, simply extending the input sequence length does not neces…

2kLong-Context Understanding

Talk like a Graph: Encoding Graphs for Large Language Models

2023-10-06 · Bahare Fatemi, Jonathan Halcrow, Bryan Perozzi

Graphs are a powerful tool for representing and analyzing complex relationships in real-world applications such as social networks, recommender systems, and computational finance. Reasoning on graphs is essential for dra…

Recommendation Systems