paper-with-me

Papers

LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models

2024-08-29 · Jingyi Wang, Jianzhong Ju, Jian Luan, Zhidong Deng

Recent advances in large vision-language models (VLMs) typically employ vision encoders based on the Vision Transformer (ViT) architecture. The division of the images into patches by ViT results in a fragmented perception, thereby hindering the visual understanding capabilities of VLMs. In this paper, we propose an innovative enhancement to address this limitation by introducing a Scene Graph Expression (SGE) module in VLMs. This module extracts and structurally expresses the complex semantic information within images, thereby improving the foundational perception and understanding abilities of VLMs. Extensive experiments demonstrate that integrating our SGE module significantly enhances the VLM's performance in vision-language tasks, indicating its effectiveness in preserving intricate semantic details and facilitating better visual understanding.

📄 PDF Abstract BibTeX arXiv:2408.16224

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Position-Wise Feed-Forward Layer 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

2024-09-26 · Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang 외

Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos. However, the developm…

3D Question Answering (3D-QA)PositionScene Understanding

SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs

2023-08-07 · Shengzhi Li, Nima Tajbakhsh

In this work, we present SciGraphQA, a synthetic multi-turn question-answer dataset related to academic graphs. SciGraphQA is 13 times larger than ChartVQA, the previously largest chart-visual question-answering dataset.…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Expressive Scene Graph Generation Using Commonsense Knowledge Infusion for Visual Understanding and Reasoning

2022-05-31 · European Semantic Web Conference (ESWC) 2022 5 · Khan, M. Jaleed; Breslin, John G.; Curry, Edward

Scene graph generation aims to capture the semantic elements in images by modelling objects and their relationships in a structured manner, which are essential for visual understanding and reasoning tasks including image…

Common Sense ReasoningGraph GenerationImage CaptioningImage Generation+9

Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs

2026-03-26 · Yike Wu, Necva Bolucu, Stephen Wan, Dadong Wang 외 arxiv

Zero-shot referring expression comprehension (REC) aims to locate target objects in images given natural language queries without relying on task-specific training data, demanding strong visual understanding capabilities…

Natural Language QueriesReferring ExpressionScene Understanding

MPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation

2026-05-11 · Ziyi Wang, Xianping Ma, Ziyao Wang, Hongyang Zhang 외 arxiv

The multimodal fusion of images and scene captions has been extensively explored and applied in various fields. However, when dealing with complex remote sensing (RS) scenes, existing studies have predominantly concentra…

Semantic SegmentationScene Segmentation