paper-with-me

홈 › Papers

PubTables-1M: Towards comprehensive table extraction from unstructured documents

2021-09-30 · CVPR 2022 1 · Brandon Smock, Rohith Pesala, Robin Abraham

Recently, significant progress has been made applying machine learning to the problem of table structure inference and extraction from unstructured documents. However, one of the greatest challenges remains the creation of datasets with complete, unambiguous ground truth at scale. To address this, we develop a new, more comprehensive dataset for table extraction, called PubTables-1M. PubTables-1M contains nearly one million tables from scientific articles, supports multiple input modalities, and contains detailed header and location information for table structures, making it useful for a wide variety of modeling approaches. It also addresses a significant source of ground truth inconsistency observed in prior datasets called oversegmentation, using a novel canonicalization procedure. We demonstrate that these improvements lead to a significant increase in training performance and a more reliable estimate of model performance at evaluation for table structure recognition. Further, we show that transformer-based object detection models trained on PubTables-1M produce excellent results for all three tasks of detection, structure recognition, and functional analysis without the need for any special customization for these tasks. Data and code will be released at https://github.com/microsoft/table-transformer.

📄 PDF Abstract BibTeX arXiv:2110.00061

Code (2)

microsoft/table-transformer 공식 구현 pytorch
phamquiluan/table-transformer pytorch

Tasks

Articlesobject-detectionObject DetectionTable DetectionTable ExtractionTable Functional AnalysisTable Recognition

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction

2025-12-11 · Brandon Smock, Valerie Faucon-Morin, Max Sokolov, Libin Liang 외 arxiv

Table extraction (TE) is a key challenge in document understanding. Traditional approaches detect tables first, then recognize their structure. Recently, interest has surged in developing methods, such as vision-language…

Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale

2025-07-09 · Javis AI Team, Amrendra Singh, Maulik Shah, Dharshan Sampath arxiv

Extracting tables and key-value pairs from financial documents is essential for business workflows such as auditing, data analytics, and automated invoice processing. In this work, we introduce Spatial ModernBERT-a trans…

Graph Neural Networks and Representation Embedding for Table Extraction in PDF Documents

2022-08-23 · Andrea Gemelli, Emanuele Vivoli, Simone Marinai

Tables are widely used in several types of documents since they can bring important information in a structured way. In scientific papers, tables can sum up novel discoveries and summarize experimental results, making th…

Optical Character Recognition (OCR)Table Extraction

Active Learning for Cascaded Object Detection: Balancing Coverage and Uncertainty in Table Extraction Pipelines

2026-07-01 · Eliott Thomas, Mickael Coustaty, Aurelie Joseph, Gaspar Deloin 외 arxiv

Table extraction from business documents relies on a cascaded pipeline where Table Detection (TD) first localizes tables and Table Structure Recognition (TSR) then recovers their internal layout. Building task-specific t…

Image ClassificationObject DetectionTable DetectionActive Learning

POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction

2026-06-08 · Brandon Smock, Libin Liang, Max Sokolov, Amrit Ramesh 외 arxiv

Large-scale document processing requires contextually aware table extraction (TE) that is both accurate and efficient. Yet current approaches require billions of parameters, hundreds of autoregressive steps, or costly AP…