paper-with-me

Papers Zero-shot Image Retrieval

“Zero-shot Image Retrieval” 태그가 달린 논문 29편 · 필터 해제

Revisiting CLIP: Efficient Alignment of 3D MRI and Tabular Data using Domain-Specific Foundation Models

2025-01-23 · Jakob Krogh Petersen, Valdemar Licht, Mads Nielsen, Asbjørn Munk

Multi-modal models require aligned, shared embedding spaces. However, common CLIP-based approaches need large amounts of samples and do not natively support 3D or tabular data, both of which are crucial in the medical do…

Image RetrievalRetrievalzero-shot-classificationZero-shot Image Retrieval+1

CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance

2024-12-05 · Chu Myaet Thwal, Ye Lin Tun, Minh N. H. Nguyen, Eui-Nam Huh 외

Beyond the success of Contrastive Language-Image Pre-training (CLIP), recent trends mark a shift toward exploring the applicability of lightweight vision-language models for resource-constrained scenarios. These models o…

Contrastive Learningcross-modal alignmentCross-Modal RetrievalLinear evaluation+6

Piecewise-Linear Manifolds for Deep Metric Learning

2024-03-22 · Shubhang Bhatnagar, Narendra Ahuja

Unsupervised deep metric learning (UDML) focuses on learning a semantic representation space using only unlabeled data. This challenging problem requires accurately estimating the similarity between data points, which is…

Image RetrievalMetric LearningRetrievalZero-shot Image Retrieval

M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining

2024-01-29 · Qingpei Guo, Furong Xu, Hanxiao Zhang, Wang Ren 외

Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative…

GPUzero-shot-classificationZero-Shot Cross-Modal RetrievalZero-shot Image Retrieval+3

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

2023-12-21 · CVPR 2024 1 · Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su 외

The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundation models, which are also critical eleme…

Image RetrievalImage-to-Text RetrievalLanguage ModellingLarge Language Model+11

Context-I2W: Mapping Images to Context-dependent Words for Accurate Zero-Shot Composed Image Retrieval

2023-09-28 · Yuanmin Tang, Jing Yu, Keke Gai, Jiamin Zhuang 외

Different from Composed Image Retrieval task that requires expensive labels for training task-specific models, Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manip…

AttributeImage RetrievalObjectRetrieval+2

GrowCLIP: Data-aware Automatic Model Growing for Large-scale Contrastive Language-Image Pre-training

2023-08-22 · ICCV 2023 1 · Xinchi Deng, Han Shi, Runhui Huang, Changlin Li 외

Cross-modal pre-training has shown impressive performance on a wide range of downstream tasks, benefiting from massive image-text pairs collected from the Internet. In practice, online data are growing constantly, highli…

image-classificationImage ClassificationImage RetrievalImage to text+2

FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing

2023-05-27 · Zhuang Li, Yuyang Chai, Terry Yue Zhuo, Lizhen Qu 외

Textual scene graph parsing has become increasingly important in various vision-language applications, including image caption evaluation and image retrieval. However, existing scene graph parsers that convert image capt…

Graph SimilarityHuman Judgment CorrelationImage CaptioningImage Retrieval+2

Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval

2023-02-06 · CVPR 2023 1 · Kuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li 외

In Composed Image Retrieval (CIR), a user combines a query image with text to describe their intended target. Existing methods rely on supervised learning of CIR models using labeled triplets consisting of the query imag…

AttributeComposed Image Retrieval (CoIR)Image RetrievalRetrieval+2

AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities

2022-11-12 · Zhongzhi Chen, Guang Liu, Bo-Wen Zhang, Fulong Ye 외

In this work, we present a conceptually simple and effective method to train a strong bilingual/multilingual multimodal representation model. Starting from the pre-trained multimodal representation model CLIP released by…

Contrastive LearningCross-Modal RetrievalImage ClassificationImage Retrieval+9

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

2022-11-02 · An Yang, Junshu Pan, Junyang Lin, Rui Men 외

The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision-language pretraining. In this work, we construct a large-scale dataset of image-text pair…

Contrastive Learningimage-classificationImage ClassificationImage Retrieval+7

General Image Descriptors for Open World Image Retrieval using ViT CLIP

2022-10-20 · Marcos V. Conde, Ivan Aerlic, Simon Jégou

The Google Universal Image Embedding (GUIE) Challenge is one of the first competitions in multi-domain image representations in the wild, covering a wide distribution of objects: landmarks, artwork, food, etc. This is a …

Image RetrievalRetrievalZero-Shot Image ClassificationZero-shot Image Retrieval+1

ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training

2022-09-30 · Bin Shan, Weichong Yin, Yu Sun, Hao Tian 외

Recent Vision-Language Pre-trained (VLP) models based on dual encoder have attracted extensive attention from academia and industry due to their superior performance on various cross-modal tasks and high computational ef…

Computational EfficiencyContrastive LearningCross-Modal RetrievalImage Retrieval+5

FETA: Towards Specializing Foundation Models for Expert Task Applications

2022-09-08 · Amit Alfassy, Assaf Arbelle, Oshri Halimi, Sivan Harary 외

Foundation Models (FMs) have demonstrated unprecedented capabilities including zero-shot learning, high fidelity data synthesis, and out of domain generalization. However, as we show in this paper, FMs still have poor ou…

Domain GeneralizationFew-Shot LearningImage RetrievalImage-text Retrieval+7

Curriculum Learning for Data-Efficient Vision-Language Alignment

2022-07-29 · Tejas Srinivasan, Xiang Ren, Jesse Thomason

Aligning image and text encoders from scratch using contrastive learning requires large amounts of paired image-text data. We alleviate this need by aligning individually pre-trained language and vision representation mo…

Contrastive LearningImage RetrievalObjectRetrieval+1

Cross-lingual and Multilingual CLIP

2022-06-01 · LREC 2022 6 · Fredrik Carlsson, Philipp Eisen, Faton Rekathati, Magnus Sahlgren

The long-standing endeavor of relating the textual and the visual domain recently underwent a pivotal breakthrough, as OpenAI released CLIP. This model distinguishes how well an English text corresponds with a given imag…

Contrastive LearningImage-text RetrievalMachine TranslationRetrieval+2

CCMB: A Large-scale Chinese Cross-modal Benchmark

2022-05-08 · Chunyu Xie, Heng Cai, Jincheng Li, Fanjing Kong 외

Vision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets a…

image-classificationImage ClassificationImage GenerationImage Retrieval+9

Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark

2022-02-14 · Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou 외

Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale dat…

BenchmarkingContrastive Learningimage-classificationImage Classification+6

Visual Representation Learning with Self-Supervised Attention for Low-Label High-data Regime

2022-01-22 · Prarthana Bhattacharyya, Chenge Li, Xiaonan Zhao, István Fehérvári 외

Self-supervision has shown outstanding results for natural language processing, and more recently, for image recognition. Simultaneously, vision transformers and its variants have emerged as a promising and scalable alte…

Few-Shot Image Classificationimage-classificationImage ClassificationImage Retrieval+4

FLAVA: A Foundational Language And Vision Alignment Model

2021-12-08 · CVPR 2022 1 · Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 외

State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal…

Image RetrievalImage-to-Text RetrievalVisual ReasoningZero-shot Image Retrieval+2
1–20 / 29 다음 →