paper-with-me

Papers

StructuralLM: Structural Pre-training for Form Understanding

2021-05-24 · ACL 2021 5 · Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, Luo Si

Large pre-trained language models achieve state-of-the-art results when fine-tuned on downstream NLP tasks. However, they almost exclusively focus on text-only representation, while neglecting cell-level layout information that is important for form image understanding. In this paper, we propose a new pre-training approach, StructuralLM, to jointly leverage cell and layout information from scanned documents. Specifically, we pre-train StructuralLM with two new designs to make the most of the interactions of cell and layout information: 1) each cell as a semantic unit; 2) classification of cell positions. The pre-trained StructuralLM achieves new state-of-the-art results in different types of downstream tasks, including form understanding (from 78.95 to 85.14), document visual question answering (from 72.59 to 83.94) and document image classification (from 94.43 to 96.08).

📄 PDF Abstract BibTeX arXiv:2105.11210

Code (1)

alibaba/AliceMind 공식 구현 pytorch

Tasks

document-image-classificationDocument Image ClassificationFormimage-classificationImage ClassificationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

SLASH the Sink: Sharpening Structural Attention Inside LLMs

2026-05-11 · Yiming Liu, Bin Lu, Xinbing Wang, Chenghu Zhou 외 arxiv

Large Language Models (LLMs) show remarkable semantic understanding but often struggle with structural understanding when processing graph topologies in a serialized format. Existing solutions rely on training external g…

ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representation

2023-11-22 · Yangyi Chen, Xingyao Wang, Manling Li, Derek Hoiem 외

State-of-the-art vision-language models (VLMs) still have limited performance in structural knowledge extraction, such as relations between objects. In this work, we present ViStruct, a training framework to learn VLMs f…

On the Emergence and Test-Time Use of Structural Information in Large Language Models

2026-01-25 · Michelle Chao Chen, Moritz Miller, Bernhard Schölkopf, Siyuan Guo arxiv

Learning structural information from observational data is central to producing new knowledge outside the training corpus. This holds for mechanistic understanding in scientific discovery as well as flexible test-time co…

Effect of structure-based training on 3D localization precision and quality

2023-09-29 · Armin Abdehkakha, Craig Snoeyink

This study introduces a structural-based training approach for CNN-based algorithms in single-molecule localization microscopy (SMLM) and 3D object reconstruction. We compare this approach with the traditional random-bas…

3D Object ReconstructionObject ReconstructionSuper-Resolution

Structural Transfer Learning in NL-to-Bash Semantic Parsers

2023-07-31 · Kyle Duffy, Satwik Bhattamishra, Phil Blunsom

Large-scale pre-training has made progress in many fields of natural language processing, though little is understood about the design of pre-training datasets. We propose a methodology for obtaining a quantitative under…

Machine TranslationSemantic ParsingTransfer LearningTranslation