Papers Prompt Learning
“Prompt Learning” 태그가 달린 논문 678편 · 필터 해제
Hierarchical Cross-modal Prompt Learning for Vision-Language Models
Pre-trained Vision-Language Models (VLMs) such as CLIP have shown excellent generalization abilities. However, adapting these large-scale models to downstream tasks while preserving their generalization capabilities rema…
Prompt LearningMGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
Recent studies have utilized visual large language models (VLMs) to answer not only "Is this face a forgery?" but also "Why is the face a forgery?" These studies introduced forgery-related attributes, such as forgery loc…
AttributeFace SwappingPrompt LearningVisual Question Answering (VQA)LifelongPR: Lifelong knowledge fusion for point cloud place recognition based on replay and prompt learning
Point cloud place recognition (PCPR) plays a crucial role in photogrammetry and robotics applications such as autonomous driving, intelligent transportation, and augmented reality. In real-world large-scale deployments o…
Autonomous DrivingContinual LearningPrompt LearningA Survey on Prompt Tuning
This survey reviews prompt tuning, a parameter-efficient approach for adapting language models by prepending trainable continuous vectors while keeping the model frozen. We classify existing approaches into two categorie…
Computational EfficiencyMixture-of-ExpertsPrompt LearningSurvey+1Integrated Structural Prompt Learning for Vision-Language Models
Prompt learning methods have significantly extended the transferability of pre-trained Vision-Language Models (VLMs) like CLIP for various downstream tasks. These methods adopt handcraft templates or learnable vectors to…
Domain GeneralizationPrompt LearningSample ProbingFA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution Detection
Pre-trained vision-language models (VLMs) have advanced out-of-distribution (OOD) detection recently. However, existing CLIP-based methods often focus on learning OOD-related knowledge to improve OOD detection, showing l…
Out-of-Distribution DetectionOut of Distribution (OOD) DetectionPrompt LearningSemantic Similarity+1Visual and Memory Dual Adapter for Multi-Modal Object Tracking
Prompt-learning-based multi-modal trackers have achieved promising progress by employing lightweight visual adapters to incorporate auxiliary modality features into frozen foundation models. However, existing approaches …
Object TrackingPrompt LearningMultimodal Prompt Alignment for Facial Expression Recognition
Prompt learning has been widely adopted to efficiently adapt vision-language models (VLMs) like CLIP for various downstream tasks. Despite their success, current VLM-based facial expression recognition (FER) methods stru…
Facial Expression RecognitionFacial Expression Recognition (FER)Large Language ModelPrompt LearningDiMPLe -- Disentangled Multi-Modal Prompt Learning: Enhancing Out-Of-Distribution Alignment with Invariant and Spurious Feature Separation
We introduce DiMPLe (Disentangled Multi-Modal Prompt Learning), a novel approach to disentangle invariant and spurious features across vision and language modalities in multi-modal learning. Spurious correlations in visu…
Contrastive LearningPrompt LearningPersonalized Federated Learning via Dual-Prompt Optimization and Cross Fusion
Federated learning (FL) enables collaborative model training across decentralized clients without sharing local data, but is challenged by heterogeneity in data, computation, and communication. Pretrained vision-language…
Federated LearningPersonalized Federated LearningPrompt LearningTaming Vision-Language Models for Medical Image Analysis: A Comprehensive Review
Modern Vision-Language Models (VLMs) exhibit unprecedented capabilities in cross-modal semantic understanding between visual and textual modalities. Given the intrinsic need for multi-modal integration in clinical applic…
Medical Image AnalysisPrompt LearningMorphSAM: Learning the Morphological Prompts from Atlases for Spine Image Segmentation
Spine image segmentation is crucial for clinical diagnosis and treatment of spine diseases. The complex structure of the spine and the high morphological similarity between individual vertebrae and adjacent intervertebra…
Image SegmentationPrompt LearningSegmentationSemantic SegmentationPrompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse…
Expressive Speech SynthesisPrompt LearningSpeech Synthesistext-to-speech+1An Empirical Study of Federated Prompt Learning for Vision Language Model
The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream tasks. However, the application of prompt …
Federated LearningLanguage ModelingLanguage ModellingPrivacy Preserving+1Foundation Molecular Grammar: Multi-Modal Foundation Models Induce Interpretable Molecular Graph Languages
Recent data-efficient molecular generation approaches exploit graph grammars to introduce interpretability into the generative models. However, grammar learning therein relies on expert annotation or unreliable heuristic…
DiversityPrompt LearningProperty PredictionValueSim: Generating Backstories to Model Individual Value Systems
As Large Language Models (LLMs) continue to exhibit increasingly human-like capabilities, aligning them with human values has become critically important. Contemporary advanced techniques, such as prompt learning and rei…
modelPrompt LearningRetrieval-augmented GenerationDiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models
Prompt learning has emerged as a powerful paradigm for adapting vision-language models such as CLIP to downstream tasks. However, existing methods often overfit to seen data, leading to significant performance degradatio…
cross-modal alignmentDomain GeneralizationFew-Shot Learningimage-classification+2SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback
Prompt quality plays a critical role in the performance of large language models (LLMs), motivating a growing body of work on prompt optimization. Most existing methods optimize prompts over a fixed dataset, assuming sta…
Prompt LearningQuestion AnsweringSynthetic Data GenerationSemantic-enhanced Co-attention Prompt Learning for Non-overlapping Cross-Domain Recommendation
Non-overlapping Cross-domain Sequential Recommendation (NCSR) is the task that focuses on domain knowledge transfer without overlapping entities. Compared with traditional Cross-domain Sequential Recommendation (CSR), NC…
Prompt LearningSequential RecommendationTransfer LearningICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification
Multi-spectral object re-identification (ReID) brings a new perception perspective for smart city and intelligent transportation applications, effectively addressing challenges from complex illumination and adverse weath…
cross-modal alignmentPrompt Learning