paper-with-me

홈 › Papers

Aligning benchmark datasets for table structure recognition

2023-03-01 · Brandon Smock, Rohith Pesala, Robin Abraham

Benchmark datasets for table structure recognition (TSR) must be carefully processed to ensure they are annotated consistently. However, even if a dataset's annotations are self-consistent, there may be significant inconsistency across datasets, which can harm the performance of models trained and evaluated on them. In this work, we show that aligning these benchmarks$\unicode{x2014}$removing both errors and inconsistency between them$\unicode{x2014}$improves model performance significantly. We demonstrate this through a data-centric approach where we adopt one model architecture, the Table Transformer (TATR), that we hold fixed throughout. Baseline exact match accuracy for TATR evaluated on the ICDAR-2013 benchmark is 65% when trained on PubTables-1M, 42% when trained on FinTabNet, and 69% combined. After reducing annotation mistakes and inter-dataset inconsistency, performance of TATR evaluated on ICDAR-2013 increases substantially to 75% when trained on PubTables-1M, 65% when trained on FinTabNet, and 81% combined. We show through ablations over the modification steps that canonicalization of the table annotations has a significantly positive effect on performance, while other choices balance necessary trade-offs that arise when deciding a benchmark dataset's final composition. Overall we believe our work has significant implications for benchmark design for TSR and potentially other tasks as well. Dataset processing and training code will be released at https://github.com/microsoft/table-transformer.

📄 PDF Abstract BibTeX arXiv:2303.00716

Code (1)

microsoft/table-transformer 공식 구현 pytorch

Tasks

Table DetectionTable Recognition

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

A Review On Table Recognition Based On Deep Learning

2023-12-08 · Shi Jiyuan, Shi chunqi

Table recognition is using the computer to automatically understand the table, to detect the position of the table from the document or picture, and to correctly extract and identify the internal structure and content of…

Data AugmentationDeep LearningTable DetectionTable Recognition

TableSeq: Unified Generation of Structure, Content, and Layout

2026-04-17 · Laziz Hamdi, Amine Tamasna, Pascal Boisson, Thierry Paquet arxiv

We present TableSeq, an image-only, end-to-end framework for joint table structure recognition, content recognition, and cell localization. The model formulates these tasks as a single sequence-generation problem: one de…

Table Recognition

An End-to-End Multi-Task Learning Model for Image-based Table Recognition

2023-03-15 · Nam Tuan Ly, Atsuhiro Takasu

Image-based table recognition is a challenging task due to the diversity of table styles and the complexity of table structures. Most of the previous methods focus on a non-end-to-end approach which divides the problem i…

Cell DetectionDecoderDiversityMulti-Task Learning+1

Rethinking Image-based Table Recognition Using Weakly Supervised Methods

2023-03-14 · Nam Tuan Ly, Atsuhiro Takasu, Phuc Nguyen, Hideaki Takeda

Most of the previous methods for table recognition rely on training datasets containing many richly annotated table images. Detailed table image annotation, e.g., cell or text bounding box annotation, however, is costly …

DecoderTable Recognition

Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner

2024-12-30 · Yitong Zhou, Mingyue Cheng, Qingyang Mao, Qi Liu 외

Pre-trained foundation models have recently significantly progressed in structured table understanding and reasoning. However, despite advancements in areas such as table semantic understanding and table question answeri…

Question AnsweringTable Recognition