paper-with-me

홈 › Papers

DocFinQA: A Long-Context Financial Reasoning Dataset

2024-01-12 · Varshini Reddy, Rik Koncel-Kedziorski, Viet Dac Lai, Michael Krumdick, Charles Lovering, Chris Tanner

For large language models (LLMs) to be effective in the financial domain -- where each decision can have a significant impact -- it is necessary to investigate realistic tasks and data. Financial professionals often interact with documents that are hundreds of pages long, but most financial research datasets only deal with short excerpts from these documents. To address this, we introduce a long-document financial QA task. We augment 7,437 questions from the existing FinQA dataset with the full-document context, extending the average context length from under 700 words in FinQA to 123k words in DocFinQA. We conduct extensive experiments over retrieval-based QA pipelines and long-context language models. DocFinQA proves a significant challenge for even state-of-the-art systems. We also provide a case-study on the longest documents in DocFinQA and find that models particularly struggle on these documents. Addressing these challenges may have a wide reaching impact across applications where specificity and long-range contexts are critical, like gene sequences and legal document contract analysis.

📄 PDF Abstract BibTeX arXiv:2401.06915

Code (0)

등록된 구현이 없습니다.

Tasks

RetrievalSpecificity

Similar Papers 제목 키워드 기반

Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning

2025-02-18 · Jingyang Lin, Andy Wong, Tian Xia, Shenghua He 외

Recent advances in Large Language Models (LLMs) have enabled them to process increasingly longer sequences, ranging from 2K to 2M tokens and even beyond. However, simply extending the input sequence length does not neces…

2kLong-Context Understanding

Document-Level Numerical Reasoning across Single and Multiple Tables in Financial Reports

2026-04-04 · Yi-Cheng Wang, Wei-An Wang, Chu-Song Chen arxiv

Despite the strong language understanding abilities of large language models (LLMs), they still struggle with reliable question answering (QA) over long, structured documents, particularly for numerical reasoning. Financ…

Question Answering

FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents

2024-11-08 · Yilun Zhao, Yitao Long, Yuru Jiang, Chengye Wang 외

We introduce FinDVer, a comprehensive benchmark specifically designed to evaluate the explainable claim verification capabilities of LLMs in the context of understanding and analyzing long, hybrid-content financial docum…

Claim VerificationRAG

Fino1: On the Transferability of Reasoning Enhanced LLMs to Finance

2025-02-12 · Lingfei Qian, Weipeng Zhou, Yan Wang, Xueqing Peng 외

While large language models (LLMs) have shown strong general reasoning capabilities, their effectiveness in financial reasoning, which is crucial for real-world financial applications remains underexplored. In this study…

BenchmarkingLong-Context Understanding

The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data

2026-06-16 · Nick Bettencourt, Xiaowei Ding, Kay Giesecke arxiv

As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of training data for large language models (LLMs). Existing long-context corpora ar…