paper-with-me

Papers

Rethinking the Evaluation of Pre-trained Text-and-Layout Models from an Entity-Centric Perspective

2024-02-04 · Chong Zhang, Yixi Zhao, Chenshu Yuan, Yi Tu, Ya Guo, Qi Zhang

Recently developed pre-trained text-and-layout models (PTLMs) have shown remarkable success in multiple information extraction tasks on visually-rich documents. However, the prevailing evaluation pipeline may not be sufficiently robust for assessing the information extraction ability of PTLMs, due to inadequate annotations within the benchmarks. Therefore, we claim the necessary standards for an ideal benchmark to evaluate the information extraction ability of PTLMs. We then introduce EC-FUNSD, an entity-centric benckmark designed for the evaluation of semantic entity recognition and entity linking on visually-rich documents. This dataset contains diverse formats of document layouts and annotations of semantic-driven entities and their relations. Moreover, this dataset disentangles the falsely coupled annotation of segment and entity that arises from the block-level annotation of FUNSD. Experiment results demonstrate that state-of-the-art PTLMs exhibit overfitting tendencies on the prevailing benchmarks, as their performance sharply decrease when the dataset bias is removed.

📄 PDF Abstract BibTeX arXiv:2402.02379

Code (1)

chongzhangFDU/ROOR 공식 구현 pytorch

Tasks

Entity LinkingSemantic entity labeling

Similar Papers 제목 키워드 기반

Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation

2024-09-07 · Jiaxin Cheng, Zixu Zhao, Tong He, Tianjun Xiao 외

Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area with…

Image GenerationLayout-to-Image GenerationVideo Editing

DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models

2025-12-01 · Patrick Kwon, Chen Chen arxiv

Current story visualization methods tend to position subjects solely by text and face challenges in maintaining artistic consistency. To address these limitations, we introduce DreamingComics, a layout-aware story visual…

Story Visualization

ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation

2025-10-13 · Ruihang Xu, Dewei Zhou, Fan Ma, Yi Yang arxiv

Multi-instance image generation (MIG) remains a significant challenge for modern diffusion models due to key limitations in achieving precise control over object layout and preserving the identity of multiple distinct su…

Image Generation

DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation

2025-04-21 · Weijie He, Mushui Liu, Yunlong Yu, Zhao Wang 외

Compositional text-to-video generation, which requires synthesizing dynamic scenes with multiple interacting entities and precise spatial-temporal relationships, remains a critical challenge for diffusion-based models. E…

AttributeDenoisingText-to-Video GenerationVideo Alignment+1

Explicitly Representing Syntax Improves Sentence-to-layout Prediction of Unexpected Situations

2024-01-25 · Wolf Nuyts, Ruben Cartuyvels, Marie-Francine Moens

Recognizing visual entities in a natural language sentence and arranging them in a 2D spatial layout require a compositional understanding of language and space. This task of layout prediction is valuable in text-to-imag…

Image GenerationSentence