paper-with-me

홈 › Papers

Visual Attention Prompted Prediction and Learning

2023-10-12 · Yifei Zhang, Siyi Gu, Bo Pan, Guangji Bai, Meikang Qiu, Xiaofeng Yang, Liang Zhao

Visual explanation (attention)-guided learning uses not only labels but also explanations to guide model reasoning process. While visual attention-guided learning has shown promising results, it requires a large number of explanation annotations that are time-consuming to prepare. However, in many real-world situations, it is usually desired to prompt the model with visual attention without model retraining. For example, when doing AI-assisted cancer classification on a medical image, users (e.g., clinicians) can provide the AI model with visual attention prompt on which areas are indispensable and which are precluded. Despite its promising objectives, achieving visual attention-prompted prediction presents several major challenges: 1) How can the visual prompt be effectively integrated into the model's reasoning process? 2) How should the model handle samples that lack visual prompts? 3) What is the impact on the model's performance when a visual prompt is imperfect? This paper introduces a novel framework for attention-prompted prediction and learning, utilizing visual prompts to steer the model's reasoning process. To improve performance in non-prompted situations and align it with prompted scenarios, we propose a co-training approach for both non-prompted and prompted models, ensuring they share similar parameters and activations. Additionally, for instances where the visual prompt does not encompass the entire input image, we have developed innovative attention prompt refinement methods. These methods interpolate the incomplete prompts while maintaining alignment with the model's explanations. Extensive experiments on four datasets demonstrate the effectiveness of our proposed framework in enhancing predictions for samples both with and without prompt.

📄 PDF Abstract BibTeX arXiv:2310.08420

Code (1)

yifeizhangcs/visual-attention-prompt 공식 구현 pytorch

Tasks

Cancer ClassificationDecision MakingPrediction

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation

2026-05-19 · Yifan Li, Xinyu Zhou, Yunhao Ge, Yu Kong arxiv

Robotic manipulation is often specified through language instructions or task identifiers, yet cluttered environments with similar objects are better handled by spatially indicating what to move and where to place it. Ad…

Trajectory Prediction

Intra and Inter Parser-Prompted Transformers for Effective Image Restoration

2025-03-18 · Cong Wang, Jinshan Pan, Liyan Wang, Wei Wang

We propose Intra and Inter Parser-Prompted Transformers (PPTformer) that explore useful features from visual foundation models for image restoration. Specifically, PPTformer contains two parts: an Image Restoration Netwo…

DeblurringImage RestorationRain Removal

Towards visually prompted keyword localisation for zero-resource spoken languages

2022-10-12 · Leanne Nortje, Herman Kamper

Imagine being able to show a system a visual depiction of a keyword and finding spoken utterances that contain this keyword from a zero-resource speech corpus. We formalise this task and call it visually prompted keyword…

PPBoost: Progressive Prompt Boosting for Text-Driven Medical Image Segmentation

2025-11-26 · Xuchen Li, Hengrui Gu, Mohan Zhang, Qin Liu 외 arxiv

Text-prompted foundation models for medical image segmentation offer an intuitive way to delineate anatomical structures from natural language queries, but their predictions often lack spatial precision and degrade under…

Medical Image SegmentationNatural Language Queries

Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach

2024-04-17 · CVPR 2024 6 · Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal, James J. Little

The emergence of attention-based transformer models has led to their extensive use in various tasks, due to their superior generalization and transfer properties. Recent research has demonstrated that such models, when p…

DecoderGeneralized Few-Shot Semantic SegmentationSemantic SegmentationVisual Prompting