paper-with-me

홈 › Papers

Locating Tables in Scanned Documents for Reconstructing and Republishing (ICIAfS14)

2014-12-24 · Akmal Jahan Mac, Roshan G. Ragel

Pool of knowledge available to the mankind depends on the source of learning resources, which can vary from ancient printed documents to present electronic material. The rapid conversion of material available in traditional libraries to digital form needs a significant amount of work if we are to maintain the format and the look of the electronic documents as same as their printed counterparts. Most of the printed documents contain not only characters and its formatting but also some associated non text objects such as tables, charts and graphical objects. It is challenging to detect them and to concentrate on the format preservation of the contents while reproducing them. To address this issue, we propose an algorithm using local thresholds for word space and line height to locate and extract all categories of tables from scanned document images. From the experiments performed on 298 documents, we conclude that our algorithm has an overall accuracy of about 75% in detecting tables from the scanned document images. Since the algorithm does not completely depend on rule lines, it can detect all categories of tables in a range of scanned documents with different font types, styles and sizes to extract their formatting features. Moreover, the algorithm can be applied to locate tables in multi column layouts with small modification in layout analysis. Treating tables with their existing formatting features will tremendously help the reproducing of printed documents for reprinting and updating purposes.

📄 PDF Abstract BibTeX arXiv:1412.7689

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ScanBank: A Benchmark Dataset for Figure Extraction from Scanned Electronic Theses and Dissertations

2021-06-23 · Sampanna Yashwant Kahu, William A. Ingram, Edward A. Fox, Jian Wu

We focus on electronic theses and dissertations (ETDs), aiming to improve access and expand their utility, since more than 6 million are publicly available, and they constitute an important corpus to aid research and edu…

Data AugmentationTable Extraction

Document Navigability: A Need for Print-Impaired

2022-06-21 · Anukriti Kumar, Tanuja Ganu, Saikat Guha

Printed documents continue to be a challenge for blind, low-vision, and other print-disabled (BLV) individuals. In this paper, we focus on the specific problem of (in-)accessibility of internal references to citations, f…

Combining Deep Learning and Reasoning for Address Detection in Unstructured Text Documents

2022-02-07 · AAAI Workshop CLeaR 2022 2 · Matthias Engelbach, Dennis Klau, Jens Drawehn, Maximilien Kintz

Extracting information from unstructured text documents is a demanding task, since these documents can have a broad variety of different layouts and a non-trivial reading order, like it is the case for multi-column docum…

Deep Learning

Reconstructing Manual Information Extraction with DB-to-Document Backprojection: Experiments in the Life Science Domain

2020-11-01 · EMNLP (sdp) 2020 11 · Mark-Christoph Müller, Sucheta Ghosh, Maja Rey, Ulrike Wittig 외

We introduce a novel scientific document processing task for making previously inaccessible information in printed paper documents available to automatic processing. We describe our data set of scanned documents and data…

Parsing Table Structures in the Wild

2021-09-06 · ICCV 2021 10 · Rujiao Long, Wen Wang, Nan Xue, Feiyu Gao 외

This paper tackles the problem of table structure parsing (TSP) from images in the wild. In contrast to existing studies that mainly focus on parsing well-aligned tabular images with simple layouts from scanned PDF docum…

Object Detection