paper-with-me

홈 › Papers

Spatial Dual-Modality Graph Reasoning for Key Information Extraction

2021-03-26 · Hongbin Sun, Zhanghui Kuang, Xiaoyu Yue, Chenhao Lin, Wayne Zhang

Key information extraction from document images is of paramount importance in office automation. Conventional template matching based approaches fail to generalize well to document images of unseen templates, and are not robust against text recognition errors. In this paper, we propose an end-to-end Spatial Dual-Modality Graph Reasoning method (SDMG-R) to extract key information from unstructured document images. We model document images as dual-modality graphs, nodes of which encode both the visual and textual features of detected text regions, and edges of which represent the spatial relations between neighboring text regions. The key information extraction is solved by iteratively propagating messages along graph edges and reasoning the categories of graph nodes. In order to roundly evaluate our proposed method as well as boost the future research, we release a new dataset named WildReceipt, which is collected and annotated tailored for the evaluation of key information extraction from document images of unseen templates in the wild. It contains 25 key information categories, a total of about 69000 text boxes, and is about 2 times larger than the existing public datasets. Extensive experiments validate that all information including visual features, textual features and spatial relations can benefit key information extraction. It has been shown that SDMG-R can effectively extract key information from document images of unseen templates, and obtain new state-of-the-art results on the recent popular benchmark SROIE and our WildReceipt. Our code and dataset will be publicly released.

📄 PDF Abstract BibTeX arXiv:2103.14470

Code (2)

open-mmlab/mmocr 공식 구현 pytorch
PaddlePaddle/PaddleOCR paddle

Tasks

Key Information ExtractionTemplate Matching

Similar Papers 제목 키워드 기반

Towards Generalizable Surgical Activity Recognition Using Spatial Temporal Graph Convolutional Networks

2020-01-11 · Duygu Sarikaya, Pierre Jannin

Modeling and recognition of surgical activities poses an interesting research problem. Although a number of recent works studied automatic recognition of surgical activities, generalizability of these works across differ…

Activity RecognitionGesture RecognitionSurgical Gesture Recognition

SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models

2025-11-28 · Ruosen Zhao, Zhikang Zhang, Jialei Xu, Jiahao Chang 외 arxiv

Large vision-language models (VLMs) show strong multimodal understanding but still struggle with 3D spatial reasoning, such as distance estimation, size comparison, and cross-view consistency. Existing 3D-aware methods e…

Spatial Reasoning

Two Heads are Better Than One: Hypergraph-Enhanced Graph Reasoning for Visual Event Ratiocination

2021-07-18 · International Conference on Machine Learning 2021 7 · Wenbo Zheng, Lan Yan, Chao Gou, Fei-Yue Wang

Even with a still image, humans can ratiocinate various visual cause-and-effect descriptions before, at present, and after, as well as beyond the given image. However, it is challenging for models to achieve such task–th…

Visual Storytelling

Reliable Multi-Modal Object Re-Identification via Modality-Aware Graph Reasoning

2025-04-21 · Xixi Wan, Aihua Zheng, Zi Wang, Bo Jiang 외

Multi-modal data provides abundant and diverse object information, crucial for effective modal interactions in Re-Identification (ReID) tasks. However, existing approaches often overlook the quality variations in local f…

Information-Theoretic Graph Fusion with Vision-Language-Action Model for Policy Reasoning and Dual Robotic Control

2025-08-07 · Shunlei Li, Longsen Gao, Jin Wang, Chang Che 외 arxiv

Teaching robots dexterous skills from human videos remains challenging due to the reliance on low-level trajectory imitation, which fails to generalize across object types, spatial layouts, and manipulator configurations…