paper-with-me

Papers

Remodeling Semantic Relationships in Vision-Language Fine-Tuning

2025-11-11 · Xiangyang Wu, Liu Liu, Baosheng Yu, Jiayan Qiu, Zhenwei Shi arxiv

Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook this information when aligning vision and language, thus leading to suboptimal performance. Toward solving this problem, we propose a method that can improve multimodal alignment and fusion based on both semantics and relationships.Specifically, we first extract multilevel semantic features from different vision encoder to capture more visual cues of the relationships. Then, we learn to project the vision features to group related semantics, among which are more likely to have relationships. Finally, we fuse the visual features with the textual by using inheritable cross-attention, where we globally remove the redundant visual relationships by discarding visual-language feature pairs with low correlation. We evaluate our proposed method on eight foundation models and two downstream tasks, visual question answering and image captioning, and show that it outperforms all existing methods.

📄 PDF Abstract BibTeX arXiv:2511.08238

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringImage Captioning

Similar Papers 제목 키워드 기반

A role for ATP-dependent chromatin remodeling in the hierarchical cooperativity between noninteracting transcription factors

2015-09-24

Chromatin remodeling machineries are abundant and diverse in eukaryotic cells. They have been involved in a variety of situations such as histone exchange and DNA repair, but their importance in gene expression remains u…

Position

Structured-Condensed Prompt Tuning in Vision-Language Models for Fine-grained Image Recognition

2026-07-07 · Xinda Liu, Qinyu Zhang, Weiqing Min, Guohua Geng 외 arxiv

Fine-grained image recognition poses a significant challenge due to the substantial expertise and effort required for manual annotation. Vision-language models (VLMs) like CLIP provide a compelling zero-shot alternative,…

Fine-Grained Image Recognition

Semantic Alignment in Hyperbolic Space for Open-Vocabulary Semantic Segmentation

2026-05-09 · Hoang M. Truong, Hai Nguyen-Truong, Dang Huynh arxiv

Open-vocabulary semantic segmentation requires adapting image-level vision-language models such as CLIP to dense pixel-level prediction, which is challenging due to the mismatch between hierarchical structure and semanti…

Semantic Segmentation

Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis

2026-07-06 · Xianhao Chen, Jiarui Hu, Yuanbo Yang, Xiyu Zhang 외 arxiv

Open-vocabulary 3D scene understanding aims to segment 3D scenes beyond predefined categories by transferring semantic knowledge from vision-language models. Existing methods have advanced this task by lifting language-a…

Scene Understanding

G^3-LQ: Marrying Hyperbolic Alignment with Explicit Semantic-Geometric Modeling for 3D Visual Grounding

2024-01-01 · CVPR 2024 1 · YuAn Wang, YaLi Li, Shengjin Wang

Grounding referred objects in 3D scenes is a burgeoning vision-language task pivotal for propelling Embodied AI as it endeavors to connect the 3D physical world with free-form descriptions. Compared to the 2D counter…

3D visual groundingVisual Grounding