paper-with-me

홈 › Papers

UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

2023-08-19 · Hao Feng, Zijian Wang, Jingqun Tang, Jinghui Lu, Wengang Zhou, Houqiang Li, Can Huang

In the era of Large Language Models (LLMs), tremendous strides have been made in the field of multimodal understanding. However, existing advanced algorithms are limited to effectively utilizing the immense representation capabilities and rich world knowledge inherent to these large pre-trained models, and the beneficial connections among tasks within the context of text-rich scenarios have not been sufficiently explored. In this work, we introduce UniDoc, a novel multimodal model equipped with text detection and recognition capabilities, which are deficient in existing approaches. Moreover, UniDoc capitalizes on the beneficial interactions among tasks to enhance the performance of each individual task. To implement UniDoc, we perform unified multimodal instruct tuning on the contributed large-scale instruction following datasets. Quantitative and qualitative experimental results show that UniDoc sets state-of-the-art scores across multiple challenging benchmarks. To the best of our knowledge, this is the first large multimodal model capable of simultaneous text detection, recognition, spotting, and understanding.

📄 PDF Abstract BibTeX arXiv:2308.11592

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingText DetectionWorld Knowledge

Similar Papers 제목 키워드 기반

UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG

2025-10-04 · Xiangyu Peng, Can Qin, Zeyuan Chen, Ran Xu 외 arxiv

Multimodal retrieval-augmented Generation (MM-RAG) is a key approach for applying large language models (LLMs) and agents to real-world knowledge bases, yet current evaluations are fragmented -- focusing on either text o…

Visual Question AnsweringLogical Reasoning

UniDoc: Unified Pretraining Framework for Document Understanding

2021-12-01 · NeurIPS 2021 12 · Jiuxiang Gu, Jason Kuen, Vlad Morariu, Handong Zhao 외

Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabeled document datasets have opened up prom…

document understandingSelf-Supervised Learning

UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

2026-04-16 · Jun Wang, Shuo Tan, Zelong Sun, Tiancheng Gu 외 arxiv

Retrieval-Augmented Generation (RAG) extends Large Vision-Language Models (LVLMs) with external visual knowledge. However, existing visual RAG systems typically rely on generic retrieval signals that overlook the fine-gr…

Reinforcement Learning

Universal Adversarial Attack on Aligned Multimodal LLMs

2025-02-11 · Temurbek Rahmatullaev, Polina Druzhinina, Matvey Mikhalchuk, Andrey Kuznetsov 외

We propose a universal adversarial attack on multimodal Large Language Models (LLMs) that leverages a single optimized image to override alignment safeguards across diverse queries and even multiple models. By backpropag…

Adversarial Attack

Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models

2025-06-02 · Youze Wang, WenBo Hu, Yinpeng Dong, Jing Liu 외

Large Language Models (LLMs) have evolved into Multimodal Large Language Models (MLLMs), significantly enhancing their capabilities by integrating visual information and other types, thus aligning more closely with the n…

Safety Alignment