paper-with-me

홈 › Papers

Towards Visual Text Grounding of Multimodal Large Language Model

2025-04-07 · Ming Li, Ruiyi Zhang, Jian Chen, Jiuxiang Gu, Yufan Zhou, Franck Dernoncourt, Wanrong Zhu, Tianyi Zhou, Tong Sun

Despite the existing evolution of Multimodal Large Language Models (MLLMs), a non-neglectable limitation remains in their struggle with visual text grounding, especially in text-rich images of documents. Document images, such as scanned forms and infographics, highlight critical challenges due to their complex layouts and textual content. However, current benchmarks do not fully address these challenges, as they mostly focus on visual grounding on natural images, rather than text-rich document images. Thus, to bridge this gap, we introduce TRIG, a novel task with a newly designed instruction dataset for benchmarking and improving the Text-Rich Image Grounding capabilities of MLLMs in document question-answering. Specifically, we propose an OCR-LLM-human interaction pipeline to create 800 manually annotated question-answer pairs as a benchmark and a large-scale training set of 90$ synthetic data based on four diverse datasets. A comprehensive evaluation of various MLLMs on our proposed benchmark exposes substantial limitations in their grounding capability on text-rich images. In addition, we propose two simple and effective TRIG methods based on general instruction tuning and plug-and-play efficient embedding, respectively. By finetuning MLLMs on our synthetic dataset, they promisingly improve spatial reasoning and grounding capabilities.

📄 PDF Abstract BibTeX arXiv:2504.04974

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelOptical Character Recognition (OCR)Question AnsweringSpatial ReasoningVisual Grounding

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

Parameter-Efficient Fine-Tuning Medical Multimodal Large Language Models for Medical Visual Grounding

2024-10-31 · Jinlong He, Pengfei Li, Gang Liu, Shenjun Zhong

Multimodal Large Language Models (MLLMs) inherit the superior text understanding capabilities of LLMs and extend these capabilities to multimodal scenarios. These models achieve excellent results in the general domain of…

parameter-efficient fine-tuningVisual Grounding

A Visual Attention Grounding Neural Model for Multimodal Machine Translation

2018-08-24 · EMNLP 2018 10 · Mingyang Zhou, Runxiang Cheng, Yong Jae Lee, Zhou Yu

We introduce a novel multimodal machine translation model that utilizes parallel visual and textual information. Our model jointly optimizes the learning of a shared visual-language embedding and a translator. The model …

Machine TranslationMultimodal Machine TranslationTranslation

Visual Grounding Strategies for Text-Only Natural Language Processing

2021-03-25 · EACL (LANTERN) 2021 4 · Damien Sileo

Visual grounding is a promising path toward more robust and accurate Natural Language Processing (NLP) models. Many multimodal extensions of BERT (e.g., VideoBERT, LXMERT, VL-BERT) allow a joint modeling of texts and ima…

Image RetrievalLanguage ModelingLanguage ModellingQuestion Answering+4

Kosmos-2: Grounding Multimodal Large Language Models to the World

2023-06-26 · Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao 외

We introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.g., bounding boxes) and grounding text to the visual world. Specifically, we represent refer…

Image CaptioningIn-Context LearningLanguage ModelingLanguage Modelling+9

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

2024-11-07 · Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao 외

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with preci…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+3