UWB@FinTOC-2020 Shared Task: Financial Document Title Detection
This paper describes our system created for the Financial Document Structure Extraction Shared Task (FinTOC-2020): Title Detection. We rely on the Apache PDFBox library to extract text and all additional information e.g. font type and font size from the financial prospectuses. Our constrained system uses only the provided training data without any additional external resources. Our system is based on the Maximum Entropy classifier and various features including font type and font size. Our system achieves F1 score 81% and #1 place in the French track and F1 score 77% and #2 place among 5 participating teams in the English track.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
UWB@FinTOC-2019 Shared Task: Financial Document Title Detection
The Financial Document Structure Extraction Shared Task (FinTOC 2022)
This paper describes the FinTOC-2022 Shared Task on the structure extraction from financial documents, its participants results and their findings. This shared task was organized as part of The 4th Financial Narrative Pr…
CILAB@FinTOC-2021 Shared Task: Title Detection and Table of Content Extraction for Financial Document
The Financial Document Structure Extraction Shared task (FinToc 2020)
This paper presents the FinTOC-2020 Shared Task on structure extraction from financial documents, its participants results and their findings. This shared task was organized as part of The 1st Joint Workshop on Financial…
Taxy.io@FinTOC-2020: Multilingual Document Structure Extraction using Transfer Learning
In this paper we describe our system submitted to the FinTOC-2020 shared task on financial doc- ument structure extraction. We propose a two-step approach to identify titles in financial docu- ments and to extract their …
Transfer Learning