GLIPv3
In the open-set object detection, the alignment of visual and text features is one of the most important factors affecting the final detection performance. This paper proposed a enhanced language and vision feature fusion module, which includes a multi-level test-image cross-attention, a text-image cross-attention and an adapted deformable self-attention. Besides, we added the deep supervison in the multi-modal task training, which is effective for the alignment of visual and text features. Experimental results show that our method performs remarkably well on COCO and LVIS datasets. Specifically, our method achieves * ********
Code (0)
등록된 구현이 없습니다.
Tasks
object-detectionObject DetectionSimilar Papers 제목 키워드 기반
GLIPv2: Unifying Localization and Vision-Language Understanding
We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2…
2D Object DetectionContrastive LearningImage CaptioningInstance Segmentation+10Hyperbolic Learning with Synthetic Captions for Open-World Detection
Open-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption dat…
HallucinationNovel ConceptsObjectobject-detection+1Elastic ViTs from Pretrained Models without Retraining
Vision foundation models achieve remarkable performance but are only available in a limited set of pre-determined sizes, forcing sub-optimal deployment choices under real-world constraints. We introduce SnapViT: Single-s…
DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment
This paper presents DetCLIPv2, an efficient and scalable training framework that incorporates large-scale image-text pairs to achieve open-vocabulary object detection (OVD). Unlike previous OVD frameworks that typically …
Language Modellingobject-detectionObject DetectionOpen-vocabulary object detection+1Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance of state-of-the-art proprietary models,…
Representation Learning