paper-with-me

홈 › Papers

Low-rank Prompt Interaction for Continual Vision-Language Retrieval

2025-01-24 · Weicai Yan, Ye Wang, Wang Lin, Zirun Guo, Zhou Zhao, Tao Jin

Research on continual learning in multi-modal tasks has been receiving increasing attention. However, most existing work overlooks the explicit cross-modal and cross-task interactions. In this paper, we innovatively propose the Low-rank Prompt Interaction (LPI) to address this general problem of multi-modal understanding, which considers both cross-modal and cross-task interactions. Specifically, as for the former, we employ multi-modal correlation modules for corresponding Transformer layers. Considering that the training parameters scale to the number of layers and tasks, we propose low-rank interaction-augmented decomposition to avoid memory explosion while enhancing the cross-modal association through sharing and separating common-specific low-rank factors. In addition, due to the multi-modal semantic differences carried by the low-rank initialization, we adopt hierarchical low-rank contrastive learning to ensure training robustness. As for the latter, we initially employ a visual analysis and identify that different tasks have clear distinctions in proximity. Therefore, we introduce explicit task contrastive constraints in the prompt learning process based on task semantic distances. Experiments on two retrieval tasks show performance improvements with the introduction of a minimal number of parameters, demonstrating the effectiveness of our method. Code is available at https://github.com/Kelvin-ywc/LPI.

📄 PDF Abstract BibTeX arXiv:2501.14369

Code (1)

kelvin-ywc/lpi 공식 구현 pytorch

Tasks

Continual LearningContrastive LearningPrompt LearningRetrieval

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
ADOPT Please enter a description about the method here

Similar Papers 제목 키워드 기반

Decouple Before Interact: Multi-Modal Prompt Learning for Continual Visual Question Answering

2023-01-01 · ICCV 2023 1 · Zi Qian, Xin Wang, Xuguang Duan, Pengda Qin 외

In the real world, a desirable Visual Question Answering model is expected to provide correct answers to new questions and images in a continual setting (recognized as CL-VQA). However, existing works formulate CLVQA…

Continual LearningLanguage ModellingPrompt LearningQuestion Answering+3

Introducing Language Guidance in Prompt-based Continual Learning

2023-08-30 · ICCV 2023 1 · Muhammad Gul Zain Ali Khan, Muhammad Ferjad Naeem, Luc van Gool, Didier Stricker 외

Continual Learning aims to learn a single model on a sequence of tasks without having access to data from previous tasks. The biggest challenge in the domain still remains catastrophic forgetting: a loss in performance o…

Continual Learning

Adaptive Rank, Reduced Forgetting: Knowledge Retention in Continual Learning Vision-Language Models with Dynamic Rank-Selective LoRA

2024-12-01 · Haodong Lu, Chongyang Zhao, Jason Xue, Lina Yao 외

We investigate whether the pre-trained knowledge of vision-language models (VLMs), such as CLIP, can be retained or even enhanced during continual learning (CL) while absorbing knowledge from a data stream. Existing meth…

Continual Learning

Q-Tuning: Queue-based Prompt Tuning for Lifelong Few-shot Language Learning

2024-04-22 · Yanhui Guo, Shaoyuan Xu, Jinmiao Fu, Jia Liu 외

This paper introduces \textbf{Q-tuning}, a novel approach for continual prompt tuning that enables the lifelong learning of a pre-trained language model. When learning a new task, Q-tuning trains a task-specific prompt b…

Language ModelingLanguage ModellingLifelong learning

MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering

2025-05-26 · Xu Li, Fan Lyu

Continual Visual Question Answering (CVQA) based on pre-trained models(PTMs) has achieved promising progress by leveraging prompt tuning to enable continual multi-modal learning. However, most existing methods adopt cros…

Continual LearningQuestion AnsweringVisual Question Answering