paper-with-me

Papers

Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks

2024-10-02 · Mengzhao Jia, Wenhao Yu, Kaixin Ma, Tianqing Fang, Zhihan Zhang, Siru Ouyang, Hongming Zhang, Meng Jiang, Dong Yu

Text-rich images, where text serves as the central visual element guiding the overall understanding, are prevalent in real-world applications, such as presentation slides, scanned documents, and webpage snapshots. Tasks involving multiple text-rich images are especially challenging, as they require not only understanding the content of individual images but reasoning about inter-relationships and logical flows across multiple visual inputs. Despite the importance of these scenarios, current multimodal large language models (MLLMs) struggle to handle such tasks due to two key challenges: (1) the scarcity of high-quality instruction tuning datasets for text-rich multi-image scenarios, and (2) the difficulty in balancing image resolution with visual feature sequence length. To address these challenges, we propose Leopard, a MLLM designed specifically for handling vision-language tasks involving multiple text-rich images. First, we curated about one million high-quality multimodal instruction-tuning data, tailored to text-rich, multi-image scenarios. Second, we developed an adaptive high-resolution multi-image encoding module to dynamically optimize the allocation of visual sequence length based on the original aspect ratios and resolutions of the input images. Experiments across a wide range of benchmarks demonstrate our model's superior capabilities in text-rich, multi-image evaluations and competitive performance in general domain evaluations.

📄 PDF Abstract BibTeX arXiv:2410.01744

Code (1)

jill0001/leopard 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Learning to Few-Shot Learn Across Diverse Natural Language Classification Tasks

2019-11-10 · COLING 2020 8 · Trapit Bansal, Rishikesh Jha, Andrew McCallum

Self-supervised pre-training of transformer models has shown enormous success in improving performance on a number of downstream tasks. However, fine-tuning on a new task still requires large amounts of task-specific lab…

DiversityEntity TypingFew-Shot LearningGeneral Classification+7

Deep Learning for Leopard Individual Identification: An Adaptive Angular Margin Approach

2024-11-04 · David Colomer Matachana

Accurate identification of individual leopards across camera trap images is critical for population monitoring and ecological studies. This paper introduces a deep learning framework to distinguish between individual leo…

Deep LearningEdge DetectionOpen Set LearningTriplet

LeoPARD --- A Generic Platform for the Implementation of Higher-Order Reasoners

2015-05-07 · Max Wisniewski, Alexander Steen, Christoph Benzmüller

LeoPARD supports the implementation of knowledge representation and reasoning tools for higher-order logic(s). It combines a sophisticated data structure layer (polymorphically typed {\lambda}-calculus with nameless spin…

LEOPARD: Parallel Optimal Deep Echo State Network Prediction Improves Service Coverage for UAV-Assisted Outdoor Hotspots

2022-03-08 · IEEE Transactions on Cognitive Communications and Networking 2022 3 · Haoran Peng, Ang-Hsun Tsai, Li-Chun Wang, Zhu Han

Unmanned aerial vehicle (UAV) base stations (BSs) can help meet the dynamic traffic demand of flash mobile crowds, but user movements also pose a significant challenge on fast-tracking for avoiding service interruption. …

Bayesian OptimizationPrediction

Named Entity Recognition in Historical Italian: The Case of Giacomo Leopardi's Zibaldone

2025-05-26 · Cristian Santini, Laura Melosi, Emanuele Frontoni

The increased digitization of world's textual heritage poses significant challenges for both computer science and literary studies. Overall, there is an urgent need of computational techniques able to adapt to the challe…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER