CoPL: Contextual Prompt Learning for Vision-Language Understanding
Recent advances in multimodal learning has resulted in powerful vision-language models, whose representations are generalizable across a variety of downstream tasks. Recently, their generalization ability has been further extended by incorporating trainable prompts, borrowed from the natural language processing literature. While such prompt learning techniques have shown impressive results, we identify that these prompts are trained based on global image features which limits itself in two aspects: First, by using global features, these prompts could be focusing less on the discriminative foreground image, resulting in poor generalization to various out-of-distribution test cases. Second, existing work weights all prompts equally whereas intuitively, prompts should be reweighed according to the semantics of the image. We address these as part of our proposed Contextual Prompt Learning (CoPL) framework, capable of aligning the prompts to the localized features of the image. Our key innovations over earlier works include using local image features as part of the prompt learning process, and more crucially, learning to weight these prompts based on local features that are appropriate for the task at hand. This gives us dynamic prompts that are both aligned to local image features as well as aware of local contextual relationships. Our extensive set of experiments on a variety of standard and few-shot datasets show that our method produces substantially improved performance when compared to the current state of the art methods. We also demonstrate both few-shot and out-of-distribution performance to establish the utility of learning dynamic prompts that are aligned to local image features.
Code (0)
등록된 구현이 없습니다.
Tasks
Prompt LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Cooperative Pseudo Labeling for Unsupervised Federated Classification
Unsupervised Federated Learning (UFL) aims to collaboratively train a global model across distributed clients without sharing data or accessing label information. Previous UFL works have predominantly focused on represen…
Representation LearningFederated LearningUnveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement
Large language models (LLMs) demonstrate exceptional instruct-following ability to complete various downstream tasks. Although this impressive ability makes LLMs flexible task solvers, their performance in solving tasks …
Combinatorial OptimizationSensitivityPedestrian Intention Prediction via Vision-Language Foundation Models
Prediction of pedestrian crossing intention is a critical function in autonomous vehicles. Conventional vision-based methods of crossing intention prediction often struggle with generalizability, context understanding, a…
Autonomous VehiclesPrompt EngineeringAutonomous DrivingManhattan Scene Understanding via XSlit Imaging
A Manhattan World (MW) [3] is composed of planar surfaces and parallel lines aligned with three mutually orthogonal principal axes. Traditional MW understanding algorithms rely on geometry priors such as the vanishing po…
3D geometryScene UnderstandingQuality factor of a transmission line coupled coplanar waveguide resonator
We investigate analytically the coupling of a coplanar waveguide resonator to a coplanar waveguide feedline. Using a conformal mapping technique we obtain an expression for the characteristic mode impedances and coupling…