Daniel@FinTOC’2 Shared Task: Title Detection and Structure Extraction
We present our contributions for the 2020 FinTOC Shared Tasks: Title Detection and Table of Contents Extraction. For the Structure Extraction task, we propose an approach that combines information from multiple sources: the table of contents, the wording of the document, and lexical domain knowledge. For the title detection task, we compare surface features to character-based features on various training configurations. We show that title detection results are very sensitive to the kind of training dataset used.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Daniel@FinTOC-2019 Shared Task : TOC Extraction and Title Detection
UWB@FinTOC-2019 Shared Task: Financial Document Title Detection
CILAB@FinTOC-2021 Shared Task: Title Detection and Table of Content Extraction for Financial Document
UWB@FinTOC-2020 Shared Task: Financial Document Title Detection
This paper describes our system created for the Financial Document Structure Extraction Shared Task (FinTOC-2020): Title Detection. We rely on the Apache PDFBox library to extract text and all additional information e.g.…
ISPRAS@FinTOC-2022 Shared Task: Two-stage TOC Generation Model
This work is connected with participation in FinTOC-2022 Shared Task: “Financial Document Structure Extraction”. The competition contains two subtasks: title detection and TOC generation. We describe an approach for solv…
Vocal Bursts Valence Prediction