paper-with-me

홈 › Papers

BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents

2021-08-10 · Teakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, Sungrae Park

Key information extraction (KIE) from document images requires understanding the contextual and spatial semantics of texts in two-dimensional (2D) space. Many recent studies try to solve the task by developing pre-trained language models focusing on combining visual features from document images with texts and their layout. On the other hand, this paper tackles the problem by going back to the basic: effective combination of text and layout. Specifically, we propose a pre-trained language model, named BROS (BERT Relying On Spatiality), that encodes relative positions of texts in 2D space and learns from unlabeled documents with area-masking strategy. With this optimized training scheme for understanding texts in 2D space, BROS shows comparable or better performance compared to previous methods on four KIE benchmarks (FUNSD, SROIE*, CORD, and SciTSR) without relying on visual features. This paper also reveals two real-world challenges in KIE tasks-(1) minimizing the error from incorrect text ordering and (2) efficient learning from fewer downstream examples-and demonstrates the superiority of BROS over previous methods. Code is available at https://github.com/clovaai/bros.

📄 PDF Abstract BibTeX arXiv:2108.04539

Code (2)

clovaai/bros 공식 구현 pytorch
2024-MindSpore-1/Code2/tree/main/model-1/bros mindspore

Tasks

Key Information ExtractionLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)Relation Extraction

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Weight Decay 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

BROS: A Pre-trained Language Model for Understanding Texts in Document

2021-01-01 · Teakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang 외

Understanding document from their visual snapshots is an emerging and challenging problem that requires both advanced computer vision and NLP methods. Although the recent advance in OCR enables the accurate extraction of…

DecoderDiversityDocument Layout Analysisdocument understanding+3

Grounded Text-to-Image Synthesis with Attention Refocusing

2023-06-08 · CVPR 2024 1 · Quynh Phung, Songwei Ge, Jia-Bin Huang

Driven by the scalable diffusion models trained on large-scale datasets, text-to-image synthesis methods have shown compelling results. However, these models still fail to precisely follow the text prompt involving multi…

Image Generation

EIGEN: Expert-Informed Joint Learning Aggregation for High-Fidelity Information Extraction from Document Images

2023-11-23 · Abhishek Singh, Venkatapathy Subramanian, Ayush Maheshwari, Pradeep Narayan 외

Information Extraction (IE) from document images is challenging due to the high variability of layout formats. Deep models such as LayoutLM and BROS have been proposed to address this problem and have shown promising res…

Liver Fibrosis and NAS scoring from CT images using self-supervised learning and texture encoding

2021-03-05 · Ananya Jana, Hui Qu, Carlos D. Minacapelli, Carolyn Catalano 외

Non-alcoholic fatty liver disease (NAFLD) is one of the most common causes of chronic liver diseases (CLD) which can progress to liver cancer. The severity and treatment of NAFLD is determined by NAFLD Activity Scores (N…

Self-Supervised LearningTransfer Learning

TOAD-GAN: Coherent Style Level Generation from a Single Example

2020-08-04 · Maren Awiszus, Frederik Schubert, Bodo Rosenhahn

In this work, we present TOAD-GAN (Token-based One-shot Arbitrary Dimension Generative Adversarial Network), a novel Procedural Content Generation (PCG) algorithm that generates token-based video game levels. TOAD-GAN fo…

Generative Adversarial Network