Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model
Image segmentation is a fundamental task in computer vision, aimed at partitioning an image into semantically meaningful regions. Referring image segmentation extends this task by using natural language expressions to localize specific objects, requiring effective integration of visual and linguistic information. In this work, we propose SegVLM, a vision-language model that incorporates architectural improvements to enhance segmentation accuracy and cross-modal alignment. The model integrates squeeze-and-excitation (SE) blocks for dynamic feature recalibration, deformable convolutions for geometric adaptability, and residual connections for deep feature learning. We also introduce a novel referring-aware fusion (RAF) loss that balances region-level alignment, boundary precision, and class imbalance. Extensive experiments and ablation studies demonstrate that each component contributes to consistent performance improvements. SegVLM also shows strong generalization across diverse datasets and referring expression scenarios.
Code (0)
등록된 구현이 없습니다.
Tasks
cross-modal alignmentImage SegmentationLanguage ModelingLanguage ModellingReferring ExpressionSegmentationSemantic SegmentationSimilar Papers 제목 키워드 기반
EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text Alignment
This work presents the network architecture EVP (Enhanced Visual Perception). EVP builds on the previous work VPD which paved the way to use the Stable Diffusion network for computer vision tasks. We propose two major en…
DecoderDepth EstimationMonocular Depth EstimationReferring Expression SegmentationReferring Segmentation in Images and Videos with Cross-Modal Self-Attention Network
We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in th…
Referring ExpressionReferring Expression SegmentationSegmentationVideo Segmentation+1Bottom-Up Shift and Reasoning for Referring Image Segmentation
Referring image segmentation aims to segment the referent that is the corresponding object or stuff referred by a natural language expression in an image. Its main challenge lies in how to effectively and efficiently…
Image SegmentationSegmentationSemantic SegmentationVisual ReasoningCross-Modal Self-Attention Network for Referring Image Segmentation
We consider the problem of referring image segmentation. Given an input image and a natural language expression, the goal is to segment the object referred by the language expression in the image. Existing works in this …
Image SegmentationReferring ExpressionReferring Expression SegmentationReferring Video Object Segmentation+1Referring Image Segmentation by Generative Adversarial Learning
Referring expression is a kind of language expression being used for referring to particular objects. In this paper, we focus on the problem of image segmentation from natural language referring expressions. Existing wor…
Image SegmentationReferring ExpressionReferring Expression SegmentationSegmentation+2