paper-with-me

홈 › Papers

Financial Numeric Extreme Labelling: A Dataset and Benchmarking for XBRL Tagging

2023-06-06 · Soumya Sharma, Subhendu Khatuya, Manjunath Hegde, Afreen Shaikh. Koustuv Dasgupta, Pawan Goyal, Niloy Ganguly

The U.S. Securities and Exchange Commission (SEC) mandates all public companies to file periodic financial statements that should contain numerals annotated with a particular label from a taxonomy. In this paper, we formulate the task of automating the assignment of a label to a particular numeral span in a sentence from an extremely large label set. Towards this task, we release a dataset, Financial Numeric Extreme Labelling (FNXL), annotated with 2,794 labels. We benchmark the performance of the FNXL dataset by formulating the task as (a) a sequence labelling problem and (b) a pipeline with span extraction followed by Extreme Classification. Although the two approaches perform comparably, the pipeline solution provides a slight edge for the least frequent labels.

📄 PDF Abstract BibTeX arXiv:2306.03723

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingSentence

Similar Papers 제목 키워드 기반

Parameter-Efficient Instruction Tuning of Large Language Models For Extreme Financial Numeral Labelling

2024-05-03 · Subhendu Khatuya, Rajdeep Mukherjee, Akash Ghosh, Manjunath Hegde 외

We study the problem of automatically annotating relevant numerals (GAAP metrics) occurring in the financial documents with their corresponding XBRL tags. Different from prior works, we investigate the feasibility of sol…

FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging

2025-06-06 · Zichen Tang, Haihong E, Ziyan Ma, Haoyang He 외

We introduce FinanceReasoning, a novel benchmark designed to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. Compared to existing benchmarks, our work provi…

Benchmarking

KPI-EDGAR: A Novel Dataset and Accompanying Metric for Relation Extraction from Financial Documents

2022-10-17 · Tobias Deußer, Syed Musharraf Ali, Lars Hillebrand, Desiana Nurchalifah 외

We introduce KPI-EDGAR, a novel dataset for Joint Named Entity Recognition and Relation Extraction building on financial reports uploaded to the Electronic Data Gathering, Analysis, and Retrieval (EDGAR) system, where th…

BenchmarkingJoint Entity and Relation Extractionnamed-entity-recognitionNamed Entity Recognition+4

Text Mining of Stocktwits Data for Predicting Stock Prices

2021-03-13 · Mukul Jaggi, Priyanka Mandal, Shreya Narang, Usman Naseem 외

Stock price prediction can be made more efficient by considering the price fluctuations and understanding the sentiments of people. A limited number of models understand financial jargon or have labelled datasets concern…

Stock Price Predictiontext-classificationText Classification

Labelling unlabelled videos from scratch with multi-modal self-supervision

2020-06-24 · NeurIPS 2020 12 · Yuki M. Asano, Mandela Patrick, Christian Rupprecht, Andrea Vedaldi

A large part of the current success of deep learning lies in the effectiveness of data -- more precisely: labelled data. Yet, labelling a dataset with human annotation continues to carry high costs, especially for videos…

BenchmarkingClustering