paper-with-me

Papers Prompt Learning

“Prompt Learning” 태그가 달린 논문 678편 · 필터 해제

Hierarchical Cross-modal Prompt Learning for Vision-Language Models

2025-07-20 · Hao Zheng, Shunzhi Yang, Zhuoxin He, Jinfeng Yang 외

Pre-trained Vision-Language Models (VLMs) such as CLIP have shown excellent generalization abilities. However, adapting these large-scale models to downstream tasks while preserving their generalization capabilities rema…

Prompt Learning

MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM

2025-07-16 · Tao Chen, Jingyi Zhang, Decheng Liu, Chunlei Peng

Recent studies have utilized visual large language models (VLMs) to answer not only "Is this face a forgery?" but also "Why is the face a forgery?" These studies introduced forgery-related attributes, such as forgery loc…

AttributeFace SwappingPrompt LearningVisual Question Answering (VQA)

LifelongPR: Lifelong knowledge fusion for point cloud place recognition based on replay and prompt learning

2025-07-14 · Xianghong Zou, Jianping Li, Zhe Chen, Zhen Cao 외

Point cloud place recognition (PCPR) plays a crucial role in photogrammetry and robotics applications such as autonomous driving, intelligent transportation, and augmented reality. In real-world large-scale deployments o…

Autonomous DrivingContinual LearningPrompt Learning

A Survey on Prompt Tuning

2025-07-08 · Zongqian Li, Yixuan Su, Nigel Collier

This survey reviews prompt tuning, a parameter-efficient approach for adapting language models by prepending trainable continuous vectors while keeping the model frozen. We classify existing approaches into two categorie…

Computational EfficiencyMixture-of-ExpertsPrompt LearningSurvey+1

Integrated Structural Prompt Learning for Vision-Language Models

2025-07-08 · Jiahui Wang, Qin Xu, Bo Jiang, Bin Luo

Prompt learning methods have significantly extended the transferability of pre-trained Vision-Language Models (VLMs) like CLIP for various downstream tasks. These methods adopt handcraft templates or learnable vectors to…

Domain GeneralizationPrompt LearningSample Probing

FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution Detection

2025-07-06 · Xinhua Lu, Runhe Lai, Yanqi Wu, Kanghao Chen 외

Pre-trained vision-language models (VLMs) have advanced out-of-distribution (OOD) detection recently. However, existing CLIP-based methods often focus on learning OOD-related knowledge to improve OOD detection, showing l…

Out-of-Distribution DetectionOut of Distribution (OOD) DetectionPrompt LearningSemantic Similarity+1

Visual and Memory Dual Adapter for Multi-Modal Object Tracking

2025-06-30 · Boyue Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu

Prompt-learning-based multi-modal trackers have achieved promising progress by employing lightweight visual adapters to incorporate auxiliary modality features into frozen foundation models. However, existing approaches …

Object TrackingPrompt Learning

Multimodal Prompt Alignment for Facial Expression Recognition

2025-06-26 · Fuyan Ma, Yiran He, Bin Sun, Shutao Li

Prompt learning has been widely adopted to efficiently adapt vision-language models (VLMs) like CLIP for various downstream tasks. Despite their success, current VLM-based facial expression recognition (FER) methods stru…

Facial Expression RecognitionFacial Expression Recognition (FER)Large Language ModelPrompt Learning

DiMPLe -- Disentangled Multi-Modal Prompt Learning: Enhancing Out-Of-Distribution Alignment with Invariant and Spurious Feature Separation

2025-06-26 · Umaima Rahman, Mohammad Yaqub, Dwarikanath Mahapatra

We introduce DiMPLe (Disentangled Multi-Modal Prompt Learning), a novel approach to disentangle invariant and spurious features across vision and language modalities in multi-modal learning. Spurious correlations in visu…

Contrastive LearningPrompt Learning

Personalized Federated Learning via Dual-Prompt Optimization and Cross Fusion

2025-06-26 · Yuguang Zhang, Kuangpu Guo, Zhihe Lu, Yunbo Wang 외

Federated learning (FL) enables collaborative model training across decentralized clients without sharing local data, but is challenged by heterogeneity in data, computation, and communication. Pretrained vision-language…

Federated LearningPersonalized Federated LearningPrompt Learning

Taming Vision-Language Models for Medical Image Analysis: A Comprehensive Review

2025-06-23 · Haoneng Lin, Cheng Xu, Jing Qin

Modern Vision-Language Models (VLMs) exhibit unprecedented capabilities in cross-modal semantic understanding between visual and textual modalities. Given the intrinsic need for multi-modal integration in clinical applic…

Medical Image AnalysisPrompt Learning

MorphSAM: Learning the Morphological Prompts from Atlases for Spine Image Segmentation

2025-06-16 · Dingwei Fan, Junyong Zhao, Chunlin Li, Xinlong Wang 외

Spine image segmentation is crucial for clinical diagnosis and treatment of spine diseases. The complex structure of the spine and the high morphological similarity between individual vertebrae and adjacent intervertebra…

Image SegmentationPrompt LearningSegmentationSemantic Segmentation

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions

2025-06-03 · Xiaoxue Gao, Huayun Zhang, Nancy F. Chen

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse…

Expressive Speech SynthesisPrompt LearningSpeech Synthesistext-to-speech+1

An Empirical Study of Federated Prompt Learning for Vision Language Model

2025-05-29 · Zhihao Wang, Wenke Huang, Tian Chen, Zekun Shi 외

The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream tasks. However, the application of prompt …

Federated LearningLanguage ModelingLanguage ModellingPrivacy Preserving+1

Foundation Molecular Grammar: Multi-Modal Foundation Models Induce Interpretable Molecular Graph Languages

2025-05-29 · Michael Sun, Weize Yuan, Gang Liu, Wojciech Matusik 외

Recent data-efficient molecular generation approaches exploit graph grammars to introduce interpretability into the generative models. However, grammar learning therein relies on expert annotation or unreliable heuristic…

DiversityPrompt LearningProperty Prediction

ValueSim: Generating Backstories to Model Individual Value Systems

2025-05-28 · Bangde Du, Ziyi Ye, Zhijing Wu, Jankowska Monika 외

As Large Language Models (LLMs) continue to exhibit increasingly human-like capabilities, aligning them with human values has become critically important. Contemporary advanced techniques, such as prompt learning and rei…

modelPrompt LearningRetrieval-augmented Generation

DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models

2025-05-26 · Niloufar Alipour Talemi, Hossein Kashiani, Hossein R. Nowdeh, Fatemeh Afghah

Prompt learning has emerged as a powerful paradigm for adapting vision-language models such as CLIP to downstream tasks. However, existing methods often overfit to seen data, leading to significant performance degradatio…

cross-modal alignmentDomain GeneralizationFew-Shot Learningimage-classification+2

SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback

2025-05-26 · Yaoning Yu, Ye Yu, Kai Wei, Haojing Luo 외

Prompt quality plays a critical role in the performance of large language models (LLMs), motivating a growing body of work on prompt optimization. Most existing methods optimize prompts over a fixed dataset, assuming sta…

Prompt LearningQuestion AnsweringSynthetic Data Generation

Semantic-enhanced Co-attention Prompt Learning for Non-overlapping Cross-Domain Recommendation

2025-05-25 · Lei Guo, Chenlong Song, Feng Guo, Xiaohui Han 외

Non-overlapping Cross-domain Sequential Recommendation (NCSR) is the task that focuses on domain knowledge transfer without overlapping entities. Compared with traditional Cross-domain Sequential Recommendation (CSR), NC…

Prompt LearningSequential RecommendationTransfer Learning

ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification

2025-05-23 · Shihao Li, Chenglong Li, Aihua Zheng, Jin Tang 외

Multi-spectral object re-identification (ReID) brings a new perception perspective for smart city and intelligent transportation applications, effectively addressing challenges from complex illumination and adverse weath…

cross-modal alignmentPrompt Learning
1–20 / 678 다음 →