paper-with-me

홈 › Papers

Unlocking UML Class Diagram Understanding in Vision Language Models

2026-05-12 · Artem Naboichenko, René Peinl arxiv

Although Vision Language Models (VLMs) have seen tremendous progress across all kinds of use cases, they still fall behind in answering questions regard-ing diagrams compared to photos. Although progress has been made in the area of bar charts, line charts and other diagrams like that there is still few research concerned with other types of diagrams, e.g. in the computer science domain. Our work presents a benchmark for visual question answering based on UML class diagrams which is both challenging and manageable. We further construct a large-scale training dataset with 16.000 image-question-answer triples and show that a LoRA-based finetune easily outperforms Qwen 3.5 27B, which is a recent and well-performing VLM in many other benchmarks.

📄 PDF Abstract BibTeX arXiv:2605.11634

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models

2025-09-02 · Hiroshi Sasaki arxiv

Multimodal models, such as the Contrastive Language-Image Pre-training (CLIP) model, have demonstrated remarkable success in aligning visual and linguistic representations. However, these models exhibit limitations when …

Visual Question AnsweringContrastive LearningImage-text matching

Do Vision-Language Models Really Understand Visual Language?

2024-09-30 · Yifan Hou, Buse Giledereli, Yilei Tu, Mrinmaya Sachan

Visual language is a system of communication that conveys information through symbols, shapes, and spatial arrangements. Diagrams are a typical example of a visual language depicting complex concepts and their relationsh…

Prior Bias in Vision Language Models on UML Diagram Interpretation

2026-07-03 · Zaiyu Cheng, Khai-Nguyen Nguyen, Antonio Mastropaolo arxiv

Vision Language Models (VLMs) are increasingly applied to software engineering artifacts, especially UML class diagrams whose meaning depends on visual notation. Yet, it is unclear whether VLMs actually read such diagram…

Pseudo Contrastive Learning for Diagram Comprehension in Multimodal Models

2026-02-27 · Hiroshi Sasaki arxiv

Recent multimodal models such as Contrastive Language-Image Pre-training (CLIP) have shown remarkable ability to align visual and linguistic representations. However, domains where small visual differences carry large se…

Visual Question AnsweringContrastive LearningImage-text matching

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions

2025-02-05 · Shue Shiinoki, Ryo Koshihara, Hayato Motegi, Masumi Morishige

Diagrams play a crucial role in visually conveying complex relationships and processes within business documentation. Despite recent advances in Vision-Language Models (VLMs) for various image understanding tasks, accura…

Language ModelingLanguage Modelling