paper-with-me

홈 › Papers

SmoGVLM: A Small, Graph-enhanced Vision-Language Model

2026-04-15 · Debjyoti Mondal, Rituraj Singh, Subhadarshi Panda arxiv

Large vision-language models (VLMs) achieve strong performance on multimodal tasks but often suffer from hallucination and poor grounding in knowledge-intensive reasoning. We propose SmoGVLM, a small, graph-enhanced VLM that integrates structured knowledge with visual and textual modalities, using Graph Neural Networks. We investigate the effects of our method across a range of model sizes, from tiny (1.3B) to large (13B) models. Our results demonstrate that, when trained using our approach, a small model can achieve performance gains upto 16.24%, and surpass its larger counterparts, outperforming larger VLMs and strong fine-tuned baselines. These results highlight the potential of structured knowledge augmentation for efficient, smaller-scale multimodal reasoning systems.

📄 PDF Abstract BibTeX arXiv:2604.16517

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

2020-06-30 · Fei Yu, Jiji Tang, Weichong Yin, Yu Sun 외

We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL tries to build the detailed semantic co…

AttributePredictionReferring Expression ComprehensionSentence+1

The ADAPT Enhanced Dependency Parser at the IWPT 2020 Shared Task

2020-09-03 · WS 2020 7 · James Barry, Joachim Wagner, Jennifer Foster

We describe the ADAPT system for the 2020 IWPT Shared Task on parsing enhanced Universal Dependencies in 17 languages. We implement a pipeline approach using UDPipe and UDPipe-future to provide initial levels of annotati…

ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation

2026-03-09 · Haoyu Tong, Xiangyu Dong, Xiaoguang Ma, Haoran Zhao 외 arxiv

Existing aerial Vision-Language Navigation (VLN) methods predominantly adopt a detection-and-planning pipeline, which converts open-vocabulary detections into discrete textual scene graphs. These approaches are plagued b…

Vision-Language NavigationSpatial Reasoning

Cephalo: Multi-Modal Vision-Language Models for Bio-Inspired Materials Analysis and Design

2024-05-29 · Markus J. Buehler

We present Cephalo, a series of multimodal vision large language models (V-LLMs) designed for materials science applications, integrating visual and linguistic data for enhanced understanding. A key innovation of Cephalo…

Dataset GenerationImage to textNatural Language UnderstandingText to 3D

Hierarchical Vision Transformer Enhanced by Graph Convolutional Network for Image Classification

2026-04-18 · Haibin Jiao arxiv

Vision Transformer (ViT) has brought new breakthroughs to the field of image classification by introducing the self-attention mechanism and Graph Convolutional Networks(GCN) have been proposed and successfully applied in…

Image Classification