paper-with-me

홈 › Papers

Image Position Prediction in Multimodal Documents

2020-05-01 · LREC 2020 5 · Masayasu Muraoka, Ryosuke Kohita, Etsuko Ishii

Conventional multimodal tasks, such as caption generation and visual question answering, have allowed machines to understand an image by describing or being asked about it in natural language, often via a sentence. Datasets for these tasks contain a large number of pairs of an image and the corresponding sentence as an instance. However, a real multimodal document such as a news article or Wikipedia page consists of multiple sentences with multiple images. Such documents require an advanced skill of jointly considering the multiple texts and multiple images, beyond a single sentence and image, for the interpretation. Therefore, aiming at building a system that can understand multimodal documents, we propose a task called image position prediction (IPP). In this task, a system learns plausible positions of images in a given document. To study this task, we automatically constructed a dataset of 66K multimodal documents with 320K images from Wikipedia articles. We conducted a preliminary experiment to evaluate the performance of a current multimodal system on our task. The experimental results show that the system outperformed simple baselines while the performance is still far from human performance, which thus poses new challenges in multimodal research.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesCaption GenerationPositionPredictionQuestion AnsweringSentenceVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

2026-05-15 · Huacan Chai, Yukai Wang, Yingxuan Yang, Dan Peng 외 arxiv

Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use evidence distributed across independently originated sources. We argue…

Multimodal Reasoning

PARL: Position-Aware Relation Learning Network for Document Layout Analysis

2026-01-12 · Fuyuan Liu, Dianyu Yu, He Ren, Nayu Liu 외 arxiv

Document layout analysis aims to detect and categorize structural elements (e.g., titles, tables, figures) in scanned or digital documents. Popular methods often rely on high-quality Optical Character Recognition (OCR) t…

Document Layout Analysis

HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation

2026-08-14 · Yin Li, Ziyang Hu, Zhiyu Guo, Xiangyu Liu 외 arxiv

Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faithful evidence selection and placement. We…

SelfDoc: Self-Supervised Document Representation Learning

2021-06-07 · CVPR 2021 1 · Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu 외

We propose SelfDoc, a task-agnostic pre-training framework for document image understanding. Because documents are multimodal and are intended for sequential reading, our framework exploits the positional, textual, and v…

Representation Learning

Multimodal Pre-training Based on Graph Attention Network for Document Understanding

2022-03-25 · Zhenrong Zhang, Jiefeng Ma, Jun Du, Licheng Wang 외

Document intelligence as a relatively new research topic supports many business applications. Its main task is to automatically read, understand, and analyze documents. However, due to the diversity of formats (invoices,…

document understandingGraph AttentionSentence