paper-with-me

Papers

LayerDoc: Layer-wise Extraction of Spatial Hierarchical Structure in Visually-Rich Documents

2023-01-07 · IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2023 1 · Puneet Mathur, Rajiv Jain, Ashutosh Mehra, Jiuxiang Gu, Franck Dernoncourt, Anandhavelu N, Quan Tran, Verena Kaynig-Fittkau, Ani Nenkova, Dinesh Manocha, Vlad I. Morariu

Digital documents often contain images and scanned text. Parsing such visually-rich documents is a core task for work-flow automation, but it remains challenging since most documents do not encode explicit layout information, e.g., how characters and words are grouped into boxes and ordered into larger semantic entities. Current state-of-the-art layout extraction methods are challenged by such documents as they rely on word sequences to have correct reading order and do not exploit their hierarchical structure. We propose LayerDoc, an approach that uses visual features, textual semantics, and spatial coordinates along with constraint inference to extract the hierarchical layout structure of documents in a bottom-up layer-wise fashion. LayerDoc recursively groups smaller regions into larger semantic elements in 2D to infer complex nested hierarchies. Experiments show that our approach outperforms competitive baselines by 10-15% on three diverse datasets of forms and mobile app screen layouts for the tasks of spatial region classification, higher-order group identification, layout hierarchy extraction, reading order detection, and word grouping.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Reading Order Detection

Similar Papers 제목 키워드 기반

GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning

2026-04-14 · Zhaochen Liu, Limeng Qiao, Guanglu Wan, Tingting Jiang arxiv

Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Recent efforts mitigate this by injecting geometric features from 3D foun…

Spatial Reasoning

Towards Generalizable Deepfake Detection with Spatial-Frequency Collaborative Learning and Hierarchical Cross-Modal Fusion

2025-04-24 · Mengyu Qiao, Runze Tian, Yang Wang

The rapid evolution of deep generative models poses a critical challenge to deepfake detection, as detectors trained on forgery-specific artifacts often suffer significant performance degradation when encountering unseen…

DeepFake DetectionFace Swapping

Shallow Network Based on Depthwise Over-Parameterized Convolution for Hyperspectral Image Classification

2021-12-01 · Hongmin Gao, Member, Zhonghao Chen, Student Member 외

Recently, convolutional neural network (CNN) techniques have gained popularity as a tool for hyperspectral image classification (HSIC). To improve the feature extraction efficiency of HSIC under the condition of limited …

Computational EfficiencyHyperspectral Image Classificationimage-classificationImage Classification

Hierarchical Spatial Sum-Product Networks for Action Recognition in Still Images

2015-11-17 · Jinghua Wang, Gang Wang

Recognizing actions from still images is popularly studied recently. In this paper, we model an action class as a flexible number of spatial configurations of body parts by proposing a new spatial SPN (Sum-Product Networ…

Action RecognitionAction Recognition In Still ImagesTemporal Action Localization

Gated Multi-layer Convolutional Feature Extraction Network for Robust Pedestrian Detection

2019-10-25 · Tianrui Liu, Jun-Jie Huang, Tianhong Dai, Guangyu Ren 외

Pedestrian detection methods have been significantly improved with the development of deep convolutional neural networks. Nevertheless, robustly detecting pedestrians with a large variant on sizes and with occlusions rem…

Pedestrian Detection