paper-with-me

홈 › Papers

SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding

2024-07-03 · Weitai Kang, Gaowen Liu, Mubarak Shah, Yan Yan

Different from Object Detection, Visual Grounding deals with detecting a bounding box for each text-image pair. This one box for each text-image data provides sparse supervision signals. Although previous works achieve impressive results, their passive utilization of annotation, i.e. the sole use of the box annotation as regression ground truth, results in a suboptimal performance. In this paper, we present SegVG, a novel method transfers the box-level annotation as Segmentation signals to provide an additional pixel-level supervision for Visual Grounding. Specifically, we propose the Multi-layer Multi-task Encoder-Decoder as the target grounding stage, where we learn a regression query and multiple segmentation queries to ground the target by regression and segmentation of the box in each decoding layer, respectively. This approach allows us to iteratively exploit the annotation as signals for both box-level regression and pixel-level segmentation. Moreover, as the backbones are typically initialized by pretrained parameters learned from unimodal tasks and the queries for both regression and segmentation are static learnable embeddings, a domain discrepancy remains among these three types of features, which impairs subsequent target grounding. To mitigate this discrepancy, we introduce the Triple Alignment module, where the query, text, and vision tokens are triangularly updated to share the same space by triple attention mechanism. Extensive experiments on five widely used datasets validate our state-of-the-art (SOTA) performance.

📄 PDF Abstract BibTeX arXiv:2407.03200

Code (1)

weitaikang/segvg 공식 구현 pytorch

Tasks

object-detectionObject DetectionregressionSegmentationVisual Grounding

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images

2026-03-20 · Jinyuan Qu, Hongyang Li, Lei Zhang arxiv

3D instance segmentation methods typically rely on high-quality point clouds or posed RGB-D scans, requiring complex multi-stage processing pipelines, and are highly sensitive to reconstruction noise. While recent feed-f…

Multi-View 3D Reconstruction3D Instance SegmentationPoint Clouds

CUPre: Cross-domain Unsupervised Pre-training for Few-Shot Cell Segmentation

2023-10-06 · Weibin Liao, Xuhong LI, Qingzhong Wang, Yanwu Xu 외

While pre-training on object detection tasks, such as Common Objects in Contexts (COCO) [1], could significantly boost the performance of cell segmentation, it still consumes on massive fine-annotated cell images [2] wit…

Cell SegmentationContrastive LearningFew-shot Instance SegmentationInstance Segmentation+5

An Exploration of Target-Conditioned Segmentation Methods for Visual Object Trackers

2020-08-03 · Matteo Dunnhofer, Niki Martinel, Christian Micheloni

Visual object tracking is the problem of predicting a target object's state in a video. Generally, bounding-boxes have been used to represent states, and a surge of effort has been spent by the community to produce effic…

Object TrackingSegmentationVisual Object Tracking

Robust Visual Tracking by Segmentation

2022-03-21 · Matthieu Paul, Martin Danelljan, Christoph Mayer, Luc van Gool

Estimating the target extent poses a fundamental challenge in visual object tracking. Typically, trackers are box-centric and fully rely on a bounding box to define the target in the scene. In practice, objects often hav…

DecoderObject TrackingSegmentationSemantic Segmentation+5

Weakly Supervised Few-shot Object Segmentation using Co-Attention with Visual and Semantic Embeddings

2020-01-26 · Mennatullah Siam, Naren Doraiswamy, Boris N. Oreshkin, Hengshuai Yao 외

Significant progress has been made recently in developing few-shot object segmentation methods. Learning is shown to be successful in few-shot segmentation settings, using pixel-level, scribbles and bounding box supervis…

Few-Shot LearningObjectOne-shot visual object segmentationSegmentation+3