paper-with-me

홈 › Papers

Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding

2024-11-08 · Jaeyoo Park, Jin Young Choi, Jeonghyung Park, Bohyung Han

We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effectively handle various font sizes within document images. To address the increasing costs of considering the multi-scale visual inputs for MLLMs, we propose the Hierarchical Visual Feature Aggregation (HVFA) module, designed to reduce the number of input tokens to LLMs. Leveraging a feature pyramid with cross-attentive pooling, our approach effectively manages the trade-off between information loss and efficiency without being affected by varying document image sizes. Furthermore, we introduce a novel instruction tuning task, which facilitates the model's text-reading capability by learning to predict the relative positions of input text, eventually minimizing the risk of truncated text caused by the limited capacity of LLMs. Comprehensive experiments validate the effectiveness of our approach, demonstrating superior performance in various document understanding tasks.

📄 PDF Abstract BibTeX arXiv:2411.05254

Code (0)

등록된 구현이 없습니다.

Tasks

document understandingOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA

2026-04-15 · Yuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang 외 arxiv

Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents. Existing OCR-free methods face a trade-off between capacity and precision: …

Visual Question AnsweringSemantic RetrievalVisual Reasoning

RESIN-EDITOR: A Schema-guided Hierarchical Event Graph Visualizer and Editor

2023-12-05 · Khanh Duy Nguyen, Zixuan Zhang, Reece Suchocki, Sha Li 외

In this paper, we present RESIN-EDITOR, an interactive event graph visualizer and editor designed for analyzing complex events. Our RESIN-EDITOR system allows users to render and freely edit hierarchical event graphs ext…

FANet: Quality-Aware Feature Aggregation Network for Robust RGB-T Tracking

2018-11-24 · Yabin Zhu, Chenglong Li, Bin Luo, Jin Tang

This paper investigates how to perform robust visual tracking in adverse and challenging conditions using complementary visual and thermal infrared data (RGBT tracking). We propose a novel deep network architecture calle…

Rgb-T TrackingVisual Tracking

Multi-Vector Index Compression in Any Modality

2026-02-24 · Hanxiang Qin, Alexander Martin, Rohan Jha, Chunsheng Zuo 외 arxiv

We study efficient multi-vector retrieval for late interaction in any modality. Late interaction has emerged as a dominant paradigm for information retrieval in text, images, visual documents, and videos, but its computa…

Information Retrieval

HAMIL: Hierarchical Aggregation-Based Multi-Instance Learning for Microscopy Image Classification

2021-03-17 · Yanlun Tu, Houchao Lei, Wei Long, Yang Yang

Multi-instance learning is common for computer vision tasks, especially in biomedical image processing. Traditional methods for multi-instance learning focus on designing feature aggregation methods and multi-instance cl…

General Classificationimage-classificationImage Classification