paper-with-me

Papers

Title block detection and information extraction for enhanced building drawings search

2025-04-11 · Alessio Lombardi, Li Duan, Ahmed Elnagar, Ahmed Zaalouk, Khalid Ismail, Edlira Vakaj

The architecture, engineering, and construction (AEC) industry still heavily relies on information stored in drawings for building construction, maintenance, compliance and error checks. However, information extraction (IE) from building drawings is often time-consuming and costly, especially when dealing with historical buildings. Drawing search can be simplified by leveraging the information stored in the title block portion of the drawing, which can be seen as drawing metadata. However, title block IE can be complex especially when dealing with historical drawings which do not follow existing standards for uniformity. This work performs a comparison of existing methods for this kind of IE task, and then proposes a novel title block detection and IE pipeline which outperforms existing methods, in particular when dealing with complex, noisy historical drawings. The pipeline is obtained by combining a lightweight Convolutional Neural Network and GPT-4o, the proposed inference pipeline detects building engineering title blocks with high accuracy, and then extract structured drawing metadata from the title blocks, which can be used for drawing search, filtering and grouping. The work demonstrates high accuracy and efficiency in IE for both vector (CAD) and hand-drawn (historical) drawings. A user interface (UI) that leverages the extracted metadata for drawing search is established and deployed on real projects, which demonstrates significant time savings. Additionally, an extensible domain-expert-annotated dataset for title block detection is developed, via an efficient AEC-friendly annotation workflow that lays the foundation for future work.

📄 PDF Abstract BibTeX arXiv:2504.08645

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Taxy.io@FinTOC-2020: Multilingual Document Structure Extraction using Transfer Learning

2020-12-01 · FNP (COLING) 2020 12 · Frederic Haase, Steffen Kirchhoff

In this paper we describe our system submitted to the FinTOC-2020 shared task on financial doc- ument structure extraction. We propose a two-step approach to identify titles in financial docu- ments and to extract their …

Transfer Learning

Daniel@FinTOC’2 Shared Task: Title Detection and Structure Extraction

2020-12-01 · FNP (COLING) 2020 12 · Emmanuel Giguet, Gaël Lejeune, Jean-Baptiste Tanguy

We present our contributions for the 2020 FinTOC Shared Tasks: Title Detection and Table of Contents Extraction. For the Structure Extraction task, we propose an approach that combines information from multiple sources: …

BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction

2025-03-25 · Jan Kohút, Martin Dočekal, Michal Hradiš, Marek Vaško

Manual digitization of bibliographic metadata is time consuming and labor intensive, especially for historical and real-world archives with highly variable formatting across documents. Despite advances in machine learnin…

document understandingobject-detectionObject DetectionOptical Character Recognition (OCR)+1

Clickbait Detection with Style-aware Title Modeling and Co-attention

2020-10-01 · CCL 2020 10 · Chuhan Wu, Fangzhao Wu, Tao Qi, Yongfeng Huang

Clickbait is a form of web content designed to attract attention and entice users to click on specific hyperlinks. The detection of clickbaits is an important task for online platforms to improve the quality of web conte…

Binary ClassificationClickbait Detection

Daniel@FinTOC-2019 Shared Task : TOC Extraction and Title Detection

2019-09-01 · WS 2019 9 · Emmanuel Giguet, Ga{\"e}l Lejeune