paper-with-me

Papers

Multidomain Document Layout Understanding using Few Shot Object Detection

2018-08-22 · Pranaydeep Singh, Srikrishna Varadarajan, Ankit Narayan Singh, Muktabh Mayank Srivastava

We try to address the problem of document layout understanding using a simple algorithm which generalizes across multiple domains while training on just few examples per domain. We approach this problem via supervised object detection method and propose a methodology to overcome the requirement of large datasets. We use the concept of transfer learning by pre-training our object detector on a simple artificial (source) dataset and fine-tuning it on a tiny domain specific (target) dataset. We show that this methodology works for multiple domains with training samples as less as 10 documents. We demonstrate the effect of each component of the methodology in the end result and show the superiority of this methodology over simple object detectors.

📄 PDF Abstract BibTeX arXiv:1808.07330

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot Object DetectionObjectobject-detectionObject DetectionTransfer Learning

Similar Papers 제목 키워드 기반

BaDLAD: A Large Multi-Domain Bengali Document Layout Analysis Dataset

2023-03-09 · Md. Istiak Hossain Shihab, Md. Rakibul Hasan, Mahfuzur Rahman Emon, Syed Mobassir Hossen 외

While strides have been made in deep learning based Bengali Optical Character Recognition (OCR) in the past decade, the absence of large Document Layout Analysis (DLA) datasets has hindered the application of OCR in docu…

BenchmarkingDeep LearningDocument Layout AnalysisOptical Character Recognition+1

One-Shot Doc Snippet Detection: Powering Search in Document Beyond Text

2022-09-12 · Abhinav Java, Shripad Deshmukh, Milan Aggarwal, Surgan Jandial 외

Active consumption of digital documents has yielded scope for research in various applications, including search. Traditionally, searching within a document has been cast as a text matching problem ignoring the rich layo…

document understandingobject-detectionObject DetectionOne-Shot Object Detection+2

LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

2022-04-18 · Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 외

Self-supervised pre-training techniques have achieved remarkable progress in Document AI. Most multimodal pre-trained models use a masked language modeling objective to learn bidirectional representations on the text mod…

cross-modal alignmentDocument AIdocument-image-classificationDocument Image Classification+16

LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding

2023-05-30 · Yi Tu, Ya Guo, Huan Chen, Jinyang Tang

Visually-rich Document Understanding (VrDU) has attracted much research attention over the past years. Pre-trained models on a large number of document images with transformer-based backbones have led to significant perf…

document-image-classificationDocument Image Classificationdocument understandingimage-classification+9

Unifying Vision, Text, and Layout for Universal Document Processing

2022-12-05 · CVPR 2023 1 · Zineng Tang, ZiYi Yang, Guoxin Wang, Yuwei Fang 외

We propose Universal Document Processing (UDOP), a foundation Document AI model which unifies text, image, and layout modalities together with varied task formats, including document understanding and generation. UDOP le…

Document AIdocument understandingImage ReconstructionVisual Question Answering (VQA)