paper-with-me

홈 › Papers

UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval

2024-12-14 · Haoyu Jiang, Zhi-Qi Cheng, Gabriel Moreira, Jiawen Zhu, Jingdong Sun, Bukun Ren, Jun-Yan He, Qi Dai, Xian-Sheng Hua

Universal Cross-Domain Retrieval (UCDR) retrieves relevant images from unseen domains and classes without semantic labels, ensuring robust generalization. Existing methods commonly employ prompt tuning with pre-trained vision-language models but are inherently limited by static prompts, reducing adaptability. We propose UCDR-Adapter, which enhances pre-trained models with adapters and dynamic prompt generation through a two-phase training strategy. First, Source Adapter Learning integrates class semantics with domain-specific visual knowledge using a Learnable Textual Semantic Template and optimizes Class and Domain Prompts via momentum updates and dual loss functions for robust alignment. Second, Target Prompt Generation creates dynamic prompts by attending to masked source prompts, enabling seamless adaptation to unseen domains and classes. Unlike prior approaches, UCDR-Adapter dynamically adapts to evolving data distributions, enhancing both flexibility and generalization. During inference, only the image branch and generated prompts are used, eliminating reliance on textual inputs for highly efficient retrieval. Extensive benchmark experiments show that UCDR-Adapter consistently outperforms ProS in most cases and other state-of-the-art methods on UCDR, U(c)CDR, and U(d)CDR settings.

📄 PDF Abstract BibTeX arXiv:2412.10680

Code (1)

fine68/ucdr2024 공식 구현 pytorch

Tasks

Retrieval

Methods 이 논문이 사용한 방법론

Adapter 설명 없음

Similar Papers 제목 키워드 기반

ProS: Prompting-to-simulate Generalized knowledge for Universal Cross-Domain Retrieval

2023-12-19 · CVPR 2024 1 · Kaipeng Fang, Jingkuan Song, Lianli Gao, Pengpeng Zeng 외

The goal of Universal Cross-Domain Retrieval (UCDR) is to achieve robust performance in generalized test scenarios, wherein data may belong to strictly unknown domains and categories during training. Recently, pre-traine…

Few-Shot LearningRetrievalText RetrievalVideo-Text Retrieval

Dec-Adapter: Exploring Efficient Decoder-Side Adapter for Bridging Screen Content and Natural Image Compression

2023-01-01 · ICCV 2023 1 · Sheng Shen, Huanjing Yue, Jingyu Yang

Natural image compression has been greatly improved in the deep learning era. However, the compression performance will be heavily degraded if the pretrained encoder is directly applied on screen content image compre…

DecoderImage CompressionTransfer Learning

Efficient Adaptation of Large Vision Transformer via Adapter Re-Composing

2023-10-10 · NeurIPS 2023 11 · Wei Dong, Dawei Yan, Zhijun Lin, Peng Wang

The advent of high-capacity pre-trained models has revolutionized problem-solving in computer vision, shifting the focus from training task-specific models to adapting pre-trained models. Consequently, effectively adapti…

ARCimage-classificationImage ClassificationTransfer Learning

Prompt Tuning based Adapter for Vision-Language Model Adaption

2023-03-24 · Jingchen Sun, Jiayu Qin, Zihao Lin, Changyou Chen

Large pre-trained vision-language (VL) models have shown significant promise in adapting to various downstream tasks. However, fine-tuning the entire network is challenging due to the massive number of model parameters. …

Few-Shot Image Classificationimage-classificationImage ClassificationLanguage Modeling+1

ADAPTERMIX: Exploring the Efficacy of Mixture of Adapters for Low-Resource TTS Adaptation

2023-05-29 · Ambuj Mehrish, Abhinav Ramesh Kashyap, Li Yingting, Navonil Majumder 외

There are significant challenges for speaker adaptation in text-to-speech for languages that are not widely spoken or for speakers with accents or dialects that are not well-represented in the training data. To address t…

Speech Synthesistext-to-speechText to Speech