Papers Open Vocabulary Attribute Detection
“Open Vocabulary Attribute Detection” 태그가 달린 논문 14편 · 필터 해제
Compositional Caching for Training-free Open-vocabulary Attribute Detection
Attribute detection is crucial for many computer vision tasks, as it enables systems to describe properties such as color, texture, and material. Current approaches often rely on labor-intensive annotation processes …
AttributeOpen Vocabulary Attribute DetectionOpen-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
The evaluation of text-generative vision-language models is a challenging yet crucial endeavor. By addressing the limitations of existing Visual Question Answering (VQA) benchmarks and proposing innovative evaluation met…
Language ModelingOpen Vocabulary Attribute DetectionVisual Question AnsweringVisual Question Answering (VQA)LOWA: Localize Objects in the Wild with Attributes
We present LOWA, a novel method for localizing objects with attributes effectively in the wild. It aims to address the insufficiency of current open-vocabulary object detectors, which are limited by the lack of instance-…
AttributeObjectobject-detectionObject Detection+1BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models. This paper proposes BLIP-2, a generic and efficient pre-training strategy that bootstraps…
Generative Visual Question AnsweringImage CaptioningImage RetrievalImage to text+13OvarNet: Towards Open-vocabulary Object Attribute Recognition
In this paper, we consider the problem of simultaneously detecting objects and inferring their visual attributes in an image, even for those with no manual annotations provided at the training stage, resembling an open-v…
AttributeKnowledge DistillationObjectobject-detection+6Reproducible scaling laws for contrastive language-image learning
Scaling up neural networks has led to remarkable performance across a wide range of tasks. Moreover, performance often follows reliable scaling laws as a function of training set size, model size, and compute, which offe…
Image ClassificationOpen Vocabulary Attribute DetectionRetrievalzero-shot-classification+3Open-vocabulary Attribute Detection
Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner. Existing open-vocabulary tasks focus on object classes, whereas research on object …
AttributeLanguage ModelingLanguage ModellingObject+2Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection
Existing open-vocabulary object detectors typically enlarge their vocabulary sizes by leveraging different forms of weak supervision. This helps generalize to novel objects at inference. Two popular forms of weak-supervi…
ObjectOpen Vocabulary Attribute DetectionOpen Vocabulary Object DetectionZero-Shot Object DetectionLocalized Vision-Language Matching for Open-vocabulary Object Detection
In this work, we propose an open-vocabulary object detection method that, based on image-caption pairs, learns to detect novel object classes along with a given set of known classes. It is a two-stage training approach t…
Language ModelingLanguage ModellingObjectobject-detection+5BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furtherm…
Image CaptioningImage-text matchingImage-text RetrievalOpen Vocabulary Attribute Detection+4Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Most existing methods in vision language pre-training rely on object-centric features extracted through object detection and make fine-grained alignments between the extracted features and texts. It is challenging for th…
Cross-Modal RetrievalImage CaptioningImage RetrievalObject+7Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Large-scale vision and language representation learning has shown promising improvements on various vision-language tasks. Most existing methods employ a transformer-based multimodal encoder to jointly model visual token…
Cross-Modal RetrievalGrounded language learningImage-text matchingImage-text Retrieval+8Learning Transferable Visual Models From Natural Language Supervision
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is n…
Action RecognitionBenchmarkingFew-Shot Image Classificationgeo-localization+21Open-Vocabulary Object Detection Using Captions
Despite the remarkable accuracy of deep neural networks in object detection, they are costly to train and scale due to supervision requirements. Particularly, learning more object categories typically requires proportion…
Objectobject-detectionObject DetectionOpen Vocabulary Attribute Detection+3