paper-with-me

Papers Visual Prompting

“Visual Prompting” 태그가 달린 논문 127편 · 필터 해제

Stepwise Decomposition and Dual-stream Focus: A Novel Approach for Training-free Camouflaged Object Segmentation

2025-06-07 · Chao Yin, Hao Li, Kequan Yang, Jide Li 외

While promptable segmentation (\textit{e.g.}, SAM) has shown promise for various segmentation tasks, it still requires manual visual prompts for each object to be segmented. In contrast, task-generic promptable segmentat…

Camouflaged Object SegmentationFeature CorrelationImage CaptioningSegmentation+3

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

2025-06-04 · Yi Lu, Jiawang Cao, Yongliang Wu, Bozheng Li 외

Multi-modal Large Language Models (MLLMs) have demonstrated remarkable reasoning capability while lack explicit mechanisms for visual grounding and segmentation, creating a gap between cognitive reasoning and visual perc…

Multimodal ReasoningReasoning SegmentationSegmentationVision-Language Segmentation+2

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering

2025-05-30 · Md Intisar Chowdhury, Kittinun Aukkapinyo, Hiroshi Fujimura, Joo Ann Woo 외

In this paper, we propose a Grid-based Local and Global Area Transcription (Grid-LoGAT) system for Video Question Answering (VideoQA). The system operates in two phases. First, extracting text transcripts from video fram…

Language ModelingLanguage ModellingLarge Language ModelQuestion Answering+2

A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis

2025-05-29 · Shengyuan Liu, Boyun Zheng, WenTing Chen, Zhihao Peng 외

Endoscopic procedures are essential for diagnosing and treating internal diseases, and multi-modal large language models (MLLMs) are increasingly applied to assist in endoscopy analysis. However, current benchmarks are l…

DiagnosticVisual PromptingVisual Question Answering (VQA)

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models

2025-05-29 · Chenbin Pan, Wenbin He, Zhengzhong Tu, Liu Ren

The recent explosive interest in the reasoning capabilities of large language models, such as DeepSeek-R1, has demonstrated remarkable success through reinforcement learning-based fine-tuning frameworks, exemplified by m…

Visual Prompting

VP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation

2025-05-21 · Niccolo Avogaro, Thomas Frick, Yagmur G. Cinar, Daniel Caraballo 외

Large-scale pretrained vision backbones have transformed computer vision by providing powerful feature extractors that enable various downstream tasks, including training-free approaches like visual prompting for semanti…

parameter-efficient fine-tuningSemantic SegmentationVisual Prompting

Vision Graph Prompting via Semantic Low-Rank Decomposition

2025-05-07 · Zixiang Ai, Zichen Liu, Jiahuan Zhou

Vision GNN (ViG) demonstrates superior performance by representing images as graph structures, providing a more natural way to capture irregular semantic patterns beyond traditional grid or sequence-based representations…

parameter-efficient fine-tuningVisual Prompting

Token Coordinated Prompt Attention is Needed for Visual Prompting

2025-05-05 · Zichen Liu, Xu Zou, Gang Hua, Jiahuan Zhou

Visual prompting techniques are widely used to efficiently fine-tune pretrained Vision Transformers (ViT) by learning a small set of shared prompts for all tokens. However, existing methods overlook the unique roles of d…

DiversityVisual Prompting

Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models

2025-04-30 · Sangmin Woo, Kang Zhou, Yun Zhou, Shuai Wang 외

Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple object-based visual prompting -- overlaying visual cues (e.g., bounding b…

HallucinationObjectObject HallucinationPrompt Engineering+1

Zoomer: Adaptive Image Focus Optimization for Black-box MLLM

2025-04-30 · Jiaxu Qian, Chendong Wang, Yifan Yang, Chaoyun Zhang 외

Recent advancements in multimodal large language models (MLLMs) have broadened the scope of vision-language tasks, excelling in applications like image captioning and interactive question-answering. However, these models…

Image CaptioningObject RecognitionQuestion AnsweringVisual Prompting

RadSAM: Segmenting 3D radiological images with a 2D promptable model

2025-04-29 · Julien Khlaut, Elodie Ferreres, Daniel Tordjman, Hélène Philippe 외

Medical image segmentation is a crucial and time-consuming task in clinical care, where mask precision is extremely important. The Segment Anything Model (SAM) offers a promising approach, as it provides an interactive i…

Image SegmentationMedical Image SegmentationOrgan SegmentationSegmentation+2

Visual and textual prompts for enhancing emotion recognition in video

2025-04-24 · Zhifeng Wang, Qixuan Zhang, Peter Zhang, Wenjia Niu 외

Vision Large Language Models (VLLMs) exhibit promising potential for multi-modal understanding, yet their application to video-based emotion recognition remains limited by insufficient spatial and contextual awareness. T…

Emotion RecognitionVideo Emotion RecognitionVisual Prompting

NVSMask3D: Hard Visual Prompting with Camera Pose Interpolation for 3D Open Vocabulary Instance Segmentation

2025-04-20 · Junyuan Fang, Zihan Wang, Yejun Zhang, Shuzhe Wang 외

Vision-language models (VLMs) have demonstrated impressive zero-shot transfer capabilities in image-level visual perception tasks. However, they fall short in 3D instance-level segmentation tasks that require accurate lo…

3D Instance Segmentation3D Open-Vocabulary Instance SegmentationDescriptiveInstance Segmentation+2

Visual Prompting for One-shot Controllable Video Editing without Inversion

2025-04-19 · CVPR 2025 1 · Zhengbo Zhang, Yuxi Zhou, Duo Peng, Joo-Hwee Lim 외

One-shot controllable video editing (OCVE) is an important yet challenging task, aiming to propagate user edits that are made -- using any image editing tool -- on the first frame of a video to all subsequent frames, whi…

Video EditingVisual Prompting

EarthGPT-X: Enabling MLLMs to Flexibly and Comprehensively Understand Multi-Source Remote Sensing Imagery

2025-04-17 · Wei zhang, Miaoxin Cai, Yaqian Ning, Tong Zhang 외

Recent advances in the visual-language area have developed natural multi-modal large language models (MLLMs) for spatial reasoning through visual prompting. However, due to remote sensing (RS) imagery containing abundant…

Large Language ModelMulti-Task LearningSpatial ReasoningVisual Prompting

Prompt-Guided Attention Head Selection for Focus-Oriented Image Retrieval

2025-04-02 · Yuji Nozawa, Yu-Chieh Lin, Kazumoto Nakamura, Youyang Ng

The goal of this paper is to enhance pretrained Vision Transformer (ViT) models for focus-oriented image retrieval with visual prompting. In real-world image retrieval scenarios, both query and database images often exhi…

Image RetrievalRetrievalVisual Prompting

Is Temporal Prompting All We Need For Limited Labeled Action Recognition?

2025-04-02 · Shreyank N Gowda, Boyan Gao, Xiao Gu, Xiaobo Jin

Video understanding has shown remarkable improvements in recent years, largely dependent on the availability of large scaled labeled datasets. Recent advancements in visual-language models, especially based on contrastiv…

Action RecognitionAllComputational EfficiencyFew-Shot Learning+2

Towards Online Multi-Modal Social Interaction Understanding

2025-03-25 · Xinpeng Li, Shijian Deng, Bolin Lai, Weiguo Pian 외

Multimodal social interaction understanding (MMSI) is critical in human-robot interaction systems. In real-world scenarios, AI agents are required to provide real-time feedback. However, existing models often depend on b…

Visual Prompting

VP-NTK: Exploring the Benefits of Visual Prompting in Differentially Private Data Synthesis

2025-03-20 · Chia-Yi Hsu, Jia-You Chen, Yu-Lin Tsai, Chih-Hsun Lin 외

Differentially private (DP) synthetic data has become the de facto standard for releasing sensitive data. However, many DP generative models suffer from the low utility of synthetic data, especially for high-resolution i…

parameter-efficient fine-tuningVisual Prompting

3DAxisPrompt: Promoting the 3D Grounding and Reasoning in GPT-4o

2025-03-17 · Dingning Liu, Cheng Wang, Peng Gao, Renrui Zhang 외

Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existing studies primarily focus on logical re…

Logical ReasoningPrompt EngineeringVisual Prompting
1–20 / 127 다음 →