Papers Visual Prompting
“Visual Prompting” 태그가 달린 논문 127편 · 필터 해제
Stepwise Decomposition and Dual-stream Focus: A Novel Approach for Training-free Camouflaged Object Segmentation
While promptable segmentation (\textit{e.g.}, SAM) has shown promise for various segmentation tasks, it still requires manual visual prompts for each object to be segmented. In contrast, task-generic promptable segmentat…
Camouflaged Object SegmentationFeature CorrelationImage CaptioningSegmentation+3RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
Multi-modal Large Language Models (MLLMs) have demonstrated remarkable reasoning capability while lack explicit mechanisms for visual grounding and segmentation, creating a gap between cognitive reasoning and visual perc…
Multimodal ReasoningReasoning SegmentationSegmentationVision-Language Segmentation+2Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
In this paper, we propose a Grid-based Local and Global Area Transcription (Grid-LoGAT) system for Video Question Answering (VideoQA). The system operates in two phases. First, extracting text transcripts from video fram…
Language ModelingLanguage ModellingLarge Language ModelQuestion Answering+2A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis
Endoscopic procedures are essential for diagnosing and treating internal diseases, and multi-modal large language models (MLLMs) are increasingly applied to assist in endoscopy analysis. However, current benchmarks are l…
DiagnosticVisual PromptingVisual Question Answering (VQA)DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
The recent explosive interest in the reasoning capabilities of large language models, such as DeepSeek-R1, has demonstrated remarkable success through reinforcement learning-based fine-tuning frameworks, exemplified by m…
Visual PromptingVP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation
Large-scale pretrained vision backbones have transformed computer vision by providing powerful feature extractors that enable various downstream tasks, including training-free approaches like visual prompting for semanti…
parameter-efficient fine-tuningSemantic SegmentationVisual PromptingVision Graph Prompting via Semantic Low-Rank Decomposition
Vision GNN (ViG) demonstrates superior performance by representing images as graph structures, providing a more natural way to capture irregular semantic patterns beyond traditional grid or sequence-based representations…
parameter-efficient fine-tuningVisual PromptingToken Coordinated Prompt Attention is Needed for Visual Prompting
Visual prompting techniques are widely used to efficiently fine-tune pretrained Vision Transformers (ViT) by learning a small set of shared prompts for all tokens. However, existing methods overlook the unique roles of d…
DiversityVisual PromptingBlack-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple object-based visual prompting -- overlaying visual cues (e.g., bounding b…
HallucinationObjectObject HallucinationPrompt Engineering+1Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
Recent advancements in multimodal large language models (MLLMs) have broadened the scope of vision-language tasks, excelling in applications like image captioning and interactive question-answering. However, these models…
Image CaptioningObject RecognitionQuestion AnsweringVisual PromptingRadSAM: Segmenting 3D radiological images with a 2D promptable model
Medical image segmentation is a crucial and time-consuming task in clinical care, where mask precision is extremely important. The Segment Anything Model (SAM) offers a promising approach, as it provides an interactive i…
Image SegmentationMedical Image SegmentationOrgan SegmentationSegmentation+2Visual and textual prompts for enhancing emotion recognition in video
Vision Large Language Models (VLLMs) exhibit promising potential for multi-modal understanding, yet their application to video-based emotion recognition remains limited by insufficient spatial and contextual awareness. T…
Emotion RecognitionVideo Emotion RecognitionVisual PromptingNVSMask3D: Hard Visual Prompting with Camera Pose Interpolation for 3D Open Vocabulary Instance Segmentation
Vision-language models (VLMs) have demonstrated impressive zero-shot transfer capabilities in image-level visual perception tasks. However, they fall short in 3D instance-level segmentation tasks that require accurate lo…
3D Instance Segmentation3D Open-Vocabulary Instance SegmentationDescriptiveInstance Segmentation+2Visual Prompting for One-shot Controllable Video Editing without Inversion
One-shot controllable video editing (OCVE) is an important yet challenging task, aiming to propagate user edits that are made -- using any image editing tool -- on the first frame of a video to all subsequent frames, whi…
Video EditingVisual PromptingEarthGPT-X: Enabling MLLMs to Flexibly and Comprehensively Understand Multi-Source Remote Sensing Imagery
Recent advances in the visual-language area have developed natural multi-modal large language models (MLLMs) for spatial reasoning through visual prompting. However, due to remote sensing (RS) imagery containing abundant…
Large Language ModelMulti-Task LearningSpatial ReasoningVisual PromptingPrompt-Guided Attention Head Selection for Focus-Oriented Image Retrieval
The goal of this paper is to enhance pretrained Vision Transformer (ViT) models for focus-oriented image retrieval with visual prompting. In real-world image retrieval scenarios, both query and database images often exhi…
Image RetrievalRetrievalVisual PromptingIs Temporal Prompting All We Need For Limited Labeled Action Recognition?
Video understanding has shown remarkable improvements in recent years, largely dependent on the availability of large scaled labeled datasets. Recent advancements in visual-language models, especially based on contrastiv…
Action RecognitionAllComputational EfficiencyFew-Shot Learning+2Towards Online Multi-Modal Social Interaction Understanding
Multimodal social interaction understanding (MMSI) is critical in human-robot interaction systems. In real-world scenarios, AI agents are required to provide real-time feedback. However, existing models often depend on b…
Visual PromptingVP-NTK: Exploring the Benefits of Visual Prompting in Differentially Private Data Synthesis
Differentially private (DP) synthetic data has become the de facto standard for releasing sensitive data. However, many DP generative models suffer from the low utility of synthetic data, especially for high-resolution i…
parameter-efficient fine-tuningVisual Prompting3DAxisPrompt: Promoting the 3D Grounding and Reasoning in GPT-4o
Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existing studies primarily focus on logical re…
Logical ReasoningPrompt EngineeringVisual Prompting