FinTagging: An LLM-ready Benchmark for Extracting and Structuring Financial Information
We introduce FinTagging, the first full-scope, table-aware XBRL benchmark designed to evaluate the structured information extraction and semantic alignment capabilities of large language models (LLMs) in the context of XBRL-based financial reporting. Unlike prior benchmarks that oversimplify XBRL tagging as flat multi-class classification and focus solely on narrative text, FinTagging decomposes the XBRL tagging problem into two subtasks: FinNI for financial entity extraction and FinCL for taxonomy-driven concept alignment. It requires models to jointly extract facts and align them with the full 10k+ US-GAAP taxonomy across both unstructured text and structured tables, enabling realistic, fine-grained evaluation. We assess a diverse set of LLMs under zero-shot settings, systematically analyzing their performance on both subtasks and overall tagging accuracy. Our results reveal that, while LLMs demonstrate strong generalization in information extraction, they struggle with fine-grained concept alignment, particularly in disambiguating closely related taxonomy entries. These findings highlight the limitations of existing LLMs in fully automating XBRL tagging and underscore the need for improved semantic reasoning and schema-aware modeling to meet the demands of accurate financial disclosure. Code is available at our GitHub repository and data is at our Hugging Face repository.
Code (1)
Tasks
Concept AlignmentMulti-class ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Multimodal Multi-Speaker Merger \& Acquisition Financial Modeling: A New Task, Dataset, and Neural Baselines
Risk prediction is an essential task in financial markets. Merger and Acquisition (M{\&}A) calls provide key insights into the claims made by company executives about the restructuring of the financial firms. Extracting …
Structuring the Unstructured: A Multi-Agent System for Extracting and Querying Financial KPIs and Guidance
Extracting structured and quantitative insights from unstructured financial filings is essential in investment research, yet remains time-consuming and resource-intensive. Conventional approaches in practice rely heavily…
Natural Language QueriesRetrievalText to SQLText-To-SQLWhy Quantitative Structuring?
Quality-designed consumer products are easy to recognize. Wouldn't it be great if the quality of financial products became just as apparent? This paper is addressed to financial practitioners. It provides an informal int…
Deep Structured Feature Networks for Table Detection and Tabular Data Extraction from Scanned Financial Document Images
Automatic table detection in PDF documents has achieved a great success but tabular data extraction are still challenging due to the integrity and noise issues in detected table areas. The accurate data extraction is ext…
Optical Character RecognitionOptical Character Recognition (OCR)Table DetectionManager Characteristics and SMEs' Restructuring Decisions: In-Court vs. Out-of-Court Restructuring
This study aims to empirically investigate the impact of managers' characteristics on their choice between in-court and out-of-court restructuring. Based on the theory of upper echelons, we tested the preferences of 342 …