TabLeX: A Benchmark Dataset for Structure and Content Information Extraction from Scientific Tables
Information Extraction (IE) from the tables present in scientific articles is challenging due to complicated tabular representations and complex embedded text. This paper presents TabLeX, a large-scale benchmark dataset comprising table images generated from scientific articles. TabLeX consists of two subsets, one for table structure extraction and the other for table content extraction. Each table image is accompanied by its corresponding LATEX source code. To facilitate the development of robust table IE tools, TabLeX contains images in different aspect ratios and in a variety of fonts. Our analysis sheds light on the shortcomings of current state-of-the-art table extraction models and shows that they fail on even simple table images. Towards the end, we experiment with a transformer-based existing baseline to report performance scores. In contrast to the static benchmarks, we plan to augment this dataset with more complex and diverse tables at regular intervals.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesTable ExtractionSimilar Papers 제목 키워드 기반
Tablext: A Combined Neural Network And Heuristic Based Table Extractor
A significant portion of the data available today is found within tables. Therefore, it is necessary to use automated table extraction to obtain thorough results when data-mining. Today's popular state-of-the-art methods…
object-detectionObject DetectionOptical Character Recognition (OCR)Table ExtractionEnhancing Network Embedding with Auxiliary Information: An Explicit Matrix Factorization Perspective
Recent advances in the field of network embedding have shown the low-dimensional network representation is playing a critical role in network analysis. However, most of the existing principles of network embedding do not…
Link PredictionNetwork EmbeddingNode ClassificationBusiness Document Information Extraction: Towards Practical Benchmarks
Information extraction from semi-structured documents is crucial for frictionless business-to-business (B2B) communication. While machine learning problems related to Document Information Extraction (IE) have been studie…
Marginalized graph autoencoder for graph clustering
Graph clustering aims to discovercommunity structures in networks, the task being fundamentally challenging mainly because the topology structure and the content of the graphs are difficult to represent for clustering an…
ClusteringGraph ClusteringGraph Representation LearningRepresentation LearningA Benchmark Suite for Template Detection and Content Extraction
Template detection and content extraction are two of the main areas of information retrieval applied to the Web. They perform different analyses over the structure and content of webpages to extract some part of the docu…
Information RetrievalRetrieval