paper-with-me

홈 › Papers

Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models

2025-09-02 · Hiroshi Sasaki arxiv

Multimodal models, such as the Contrastive Language-Image Pre-training (CLIP) model, have demonstrated remarkable success in aligning visual and linguistic representations. However, these models exhibit limitations when applied to specialised visual domains, such as diagrams, which encode structured, symbolic information distinct from that of natural imagery. In this paper, we introduce a novel training paradigm explicitly designed to enhance the comprehension of diagrammatic images within vision-language models. Our approach uses ``hard'' samples for our proposed contrastive learning that incorporates two specialised loss functions that leverage the inherent structural properties of diagrams. By integrating these objectives into model training, our method enables models to develop a more structured and semantically coherent understanding of diagrammatic content. We empirically validate our approach on a benchmark dataset of flowcharts, as a representative class of diagrammatic imagery, demonstrating substantial improvements over standard CLIP and conventional hard negative CLIP learning paradigms for both image-text matching and visual question answering tasks. Our findings underscore the significance of tailored training strategies for specialised tasks and contribute to advancing diagrammatic understanding within the broader landscape of vision-language integration.

📄 PDF Abstract BibTeX arXiv:2509.01959

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringContrastive LearningImage-text matching

Similar Papers 제목 키워드 기반

Pseudo Contrastive Learning for Diagram Comprehension in Multimodal Models

2026-02-27 · Hiroshi Sasaki arxiv

Recent multimodal models such as Contrastive Language-Image Pre-training (CLIP) have shown remarkable ability to align visual and linguistic representations. However, domains where small visual differences carry large se…

Visual Question AnsweringContrastive LearningImage-text matching

Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation

2025-04-13 · Zhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li 외

Scientific diagrams are vital tools for communicating structured knowledge across disciplines. However, they are often published as static raster images, losing symbolic semantics and limiting reuse. While Multimodal Lar…

Code GenerationMultimodal Reasoningvalid

AI2D-RST: A multimodal corpus of 1000 primary school science diagrams

2019-12-09 · Tuomo Hiippala, Malihe Alikhani, Jonas Haverinen, Timo Kalliokoski 외

This article introduces AI2D-RST, a multimodal corpus of 1000 English-language diagrams that represent topics in primary school natural sciences, such as food webs, life cycles, moon phases and human physiology. The corp…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning

2026-03-31 · Taiting Lu, Kaiyuan Lin, Yuxin Tian, Mingjia Wang 외 arxiv

Recent large multimodal models (LMMs) have made rapid progress in visual grounding, document understanding, and diagram reasoning tasks. However, their ability to convert Printed Circuit Board (PCB) schematic diagrams in…

Visual GroundingVisual Reasoning

Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models

2025-05-26 · Kai Sun, Yushi Bai, Zhen Yang, Jiajie Zhang 외

Benefiting from contrastively trained visual encoders on large-scale natural scene images, Large Multimodal Models (LMMs) have achieved remarkable performance across various visual perception tasks. However, the inherent…

Contrastive LearningMath