paper-with-me

Papers

GLIPv3

2020-02-02 · CVPR 2020 2 · Jiaxing Zhao

In the open-set object detection, the alignment of visual and text features is one of the most important factors affecting the final detection performance. This paper proposed a enhanced language and vision feature fusion module, which includes a multi-level test-image cross-attention, a text-image cross-attention and an adapted deformable self-attention. Besides, we added the deep supervison in the multi-modal task training, which is effective for the alignment of visual and text features. Experimental results show that our method performs remarkably well on COCO and LVIS datasets. Specifically, our method achieves * ********

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject Detection

Similar Papers 제목 키워드 기반

GLIPv2: Unifying Localization and Vision-Language Understanding

2022-06-12 · Haotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen 외

We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2…

2D Object DetectionContrastive LearningImage CaptioningInstance Segmentation+10

Hyperbolic Learning with Synthetic Captions for Open-World Detection

2024-04-07 · CVPR 2024 1 · Fanjie Kong, Yanbei Chen, Jiarui Cai, Davide Modolo

Open-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption dat…

HallucinationNovel ConceptsObjectobject-detection+1

Elastic ViTs from Pretrained Models without Retraining

2025-10-20 · Walter Simoncini, Michael Dorkenwald, Tijmen Blankevoort, Cees G. M. Snoek 외 arxiv

Vision foundation models achieve remarkable performance but are only available in a limited set of pre-determined sizes, forcing sub-optimal deployment choices under real-world constraints. We introduce SnapViT: Single-s…

DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment

2023-04-10 · CVPR 2023 1 · Lewei Yao, Jianhua Han, Xiaodan Liang, Dan Xu 외

This paper presents DetCLIPv2, an efficient and scalable training framework that incorporates large-scale image-text pairs to achieve open-vocabulary object detection (OVD). Unlike previous OVD frameworks that typically …

Language Modellingobject-detectionObject DetectionOpen-vocabulary object detection+1

Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning

2025-07-18 · Shashanka Venkataramanan, Valentinos Pariza, Mohammadreza Salehi, Lukas Knobel 외 arxiv

We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance of state-of-the-art proprietary models,…

Representation Learning