paper-with-me

홈 › Papers

DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding

2020-10-15 · Findings of the Association for Computational Linguistics 2020 · Zilong Wang, Mingjie Zhan, Xuebo Liu, Ding Liang

Form understanding depends on both textual contents and organizational structure. Although modern OCR performs well, it is still challenging to realize general form understanding because forms are commonly used and of various formats. The table detection and handcrafted features in previous works cannot apply to all forms because of their requirements on formats. Therefore, we concentrate on the most elementary components, the key-value pairs, and adopt multimodal methods to extract features. We consider the form structure as a tree-like or graph-like hierarchy of text fragments. The parent-child relation corresponds to the key-value pairs in forms. We utilize the state-of-the-art models and design targeted extraction modules to extract multimodal features from semantic contents, layout information, and visual images. A hybrid fusion method of concatenation and feature shifting is designed to fuse the heterogeneous features and provide an informative joint representation. We adopt an asymmetric algorithm and negative sampling in our model as well. We validate our method on two benchmarks, MedForm and FUNSD, and extensive experiments demonstrate the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2010.11685

Code (0)

등록된 구현이 없습니다.

Tasks

FormOptical Character Recognition (OCR)Table Detection

Similar Papers 제목 키워드 기반

HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation

2026-08-14 · Yin Li, Ziyang Hu, Zhiyu Guo, Xiangyu Liu 외 arxiv

Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faithful evidence selection and placement. We…

mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

2024-03-19 · Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan 외

Structure information is critical for understanding the semantics of text-rich images, such as documents, tables, and charts. Existing Multimodal Large Language Models (MLLMs) for Visual Document Understanding are equipp…

document understandingOptical Character Recognition (OCR)

MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents

2026-04-14 · Joongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo 외 arxiv

RAG-based QA has emerged as a powerful method for processing long industrial documents. However, conventional text chunking approaches often neglect complex and long industrial document structures, causing information lo…

DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

2024-10-16 · Zhiyuan Zhao, Hengrui Kang, Bin Wang, Conghui He

Document Layout Analysis is crucial for real-world document understanding systems, but it encounters a challenging trade-off between speed and accuracy: multimodal methods leveraging both text and visual features achieve…

Document Layout Analysisdocument understandingModel Optimization

Extracting Variable-Depth Logical Document Hierarchy from Long Documents: Method, Evaluation, and Application

2021-05-14 · Rongyu Cao, Yixuan Cao, Ganbin Zhou, Ping Luo

In this paper, we study the problem of extracting variable-depth "logical document hierarchy" from long documents, namely organizing the recognized "physical document objects" into hierarchical structures. The discovery …

Binary ClassificationPassage RetrievalRetrieval