Open Vocabulary Attribute Detection
2개 벤치마크 · 논문 14편 · 이 태스크의 논문 보기 →
Benchmarks
OVAD-Box benchmark
OVAD benchmark
Most implemented
Learning Transferable Visual Models From Natural Language Supervision
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Reproducible scaling laws for contrastive language-image learning
Papers
Compositional Caching for Training-free Open-vocabulary Attribute Detection
Attribute detection is crucial for many computer vision tasks, as it enables systems to describe properties such as color, texture, and material. Current approaches often rely on labor-intensive annotation processes …
AttributeOpen Vocabulary Attribute DetectionOpen-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
The evaluation of text-generative vision-language models is a challenging yet crucial endeavor. By addressing the limitations of existing Visual Question Answering (VQA) benchmarks and proposing innovative evaluation met…
Language ModelingOpen Vocabulary Attribute DetectionVisual Question AnsweringVisual Question Answering (VQA)LOWA: Localize Objects in the Wild with Attributes
We present LOWA, a novel method for localizing objects with attributes effectively in the wild. It aims to address the insufficiency of current open-vocabulary object detectors, which are limited by the lack of instance-…
AttributeObjectobject-detectionObject Detection+1BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models. This paper proposes BLIP-2, a generic and efficient pre-training strategy that bootstraps…
Generative Visual Question AnsweringImage CaptioningImage RetrievalImage to text+13OvarNet: Towards Open-vocabulary Object Attribute Recognition
In this paper, we consider the problem of simultaneously detecting objects and inferring their visual attributes in an image, even for those with no manual annotations provided at the training stage, resembling an open-v…
AttributeKnowledge DistillationObjectobject-detection+6Reproducible scaling laws for contrastive language-image learning
Scaling up neural networks has led to remarkable performance across a wide range of tasks. Moreover, performance often follows reliable scaling laws as a function of training set size, model size, and compute, which offe…
Image ClassificationOpen Vocabulary Attribute DetectionRetrievalzero-shot-classification+3