paper-with-me

홈 › Papers

Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model

2025-05-25 · Alaa Dalaq, Muzammil Behzad

Image segmentation is a fundamental task in computer vision, aimed at partitioning an image into semantically meaningful regions. Referring image segmentation extends this task by using natural language expressions to localize specific objects, requiring effective integration of visual and linguistic information. In this work, we propose SegVLM, a vision-language model that incorporates architectural improvements to enhance segmentation accuracy and cross-modal alignment. The model integrates squeeze-and-excitation (SE) blocks for dynamic feature recalibration, deformable convolutions for geometric adaptability, and residual connections for deep feature learning. We also introduce a novel referring-aware fusion (RAF) loss that balances region-level alignment, boundary precision, and class imbalance. Extensive experiments and ablation studies demonstrate that each component contributes to consistent performance improvements. SegVLM also shows strong generalization across diverse datasets and referring expression scenarios.

📄 PDF Abstract BibTeX arXiv:2505.19242

Code (0)

등록된 구현이 없습니다.

Tasks

cross-modal alignmentImage SegmentationLanguage ModelingLanguage ModellingReferring ExpressionSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text Alignment

2023-12-13 · Mykola Lavreniuk, Shariq Farooq Bhat, Matthias Müller, Peter Wonka

This work presents the network architecture EVP (Enhanced Visual Perception). EVP builds on the previous work VPD which paved the way to use the Stable Diffusion network for computer vision tasks. We propose two major en…

DecoderDepth EstimationMonocular Depth EstimationReferring Expression Segmentation

Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

2021-02-09 · Linwei Ye, Mrigank Rochan, Zhi Liu, Xiaoqin Zhang 외

We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in th…

Referring ExpressionReferring Expression SegmentationSegmentationVideo Segmentation+1

Bottom-Up Shift and Reasoning for Referring Image Segmentation

2021-06-19 · CVPR 2021 1 · Sibei Yang, Meng Xia, Guanbin Li, Hong-Yu Zhou 외

Referring image segmentation aims to segment the referent that is the corresponding object or stuff referred by a natural language expression in an image. Its main challenge lies in how to effectively and efficiently…

Image SegmentationSegmentationSemantic SegmentationVisual Reasoning

Cross-Modal Self-Attention Network for Referring Image Segmentation

2019-04-09 · CVPR 2019 6 · Linwei Ye, Mrigank Rochan, Zhi Liu, Yang Wang

We consider the problem of referring image segmentation. Given an input image and a natural language expression, the goal is to segment the object referred by the language expression in the image. Existing works in this …

Image SegmentationReferring ExpressionReferring Expression SegmentationReferring Video Object Segmentation+1

Referring Image Segmentation by Generative Adversarial Learning

2020-04-20 · IEEE 2020 4 · Shuang Qiu, Yao Zhao, Jianbo Jiao, Yunchao Wei 외

Referring expression is a kind of language expression being used for referring to particular objects. In this paper, we focus on the problem of image segmentation from natural language referring expressions. Existing wor…

Image SegmentationReferring ExpressionReferring Expression SegmentationSegmentation+2