paper-with-me

홈 › Papers

Beyond Bag-of-Patches: Learning Global Layout via Textual Supervision for Late-Interaction Visual Document Retrieval

2026-05-08 · Pascal Tilli, Mohsen Mesgar arxiv

Visual Document Retrieval (VDR) models mostly rely on late interaction architectures, in which documents are represented by a set of local patch embeddings and then matched against query tokens. While efficient, this architecture prioritizes local similarity over global layout structure of documents to estimate relevancy between documents and query. In practice, this leads to errors as relevance originates from layout structure of documents with heterogeneous layouts combining figures, tables, and text. We make document layout learnable without changing inference. We propose a multimodal encoder that augments local patch representations with a global layout embedding, trained via textual descriptions encoding document layout information. Across four ViDoRe-v2 datasets, our model improves over the strongest architecturally comparable ColPali/ColQwen baseline by +2.4 nDCG@5 and +2.3 MAP@5, with statistically significant per-dataset gains over ColQwen.

📄 PDF Abstract BibTeX arXiv:2605.08421

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Any-resolution Training for High-resolution Image Synthesis

2022-04-14 · Lucy Chai, Michael Gharbi, Eli Shechtman, Phillip Isola 외

Generative models operate at fixed resolution, even though natural images come in a variety of sizes. As high-resolution details are downsampled away and low-resolution images are discarded altogether, precious supervisi…

2kImage GenerationVocal Bursts Intensity Prediction

High-Resolution Artwork Outpainting with Global Blueprint Guidance and Layout Control

2026-07-07 · Junha Kim, Hyunjoon Park, Donghyeon Cho arxiv

Image outpainting extends an image beyond its original borders, requiring seamless style integration and globally coherent scene completion. Building on the success of diffusion models, recent methods have achieved subst…

Image Outpainting

PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation

2025-08-29 · Jiho Choi, Seojeong Park, Seongjong Song, Hyunjung Shim arxiv

Automating scientific poster generation requires hierarchical document understanding and coherent content-layout planning. Existing methods often rely on flat summarization or optimize content and layout separately. As a…

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration

2026-06-07 · Jeonghwan Kim, Yushi Lan, Yongwei Chen, Hieu Trung Nguyen 외 arxiv

Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual evidence. Despite recent progress in joi…

Scene Generation

DocSynthv2: A Practical Autoregressive Modeling for Document Generation

2024-06-12 · Sanket Biswas, Rajiv Jain, Vlad I. Morariu, Jiuxiang Gu 외

While the generation of document layouts has been extensively explored, comprehensive document generation encompassing both layout and content presents a more complex challenge. This paper delves into this advanced domai…

Layout Generation