Papers Visual Prompting
“Visual Prompting” 태그가 달린 논문 127편 · 필터 해제
KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation
With the rapid advancement of large language models (LLMs) and vision-language models (VLMs), significant progress has been made in developing open-vocabulary robotic manipulation systems. However, many existing approach…
ObjectVisual PromptingChameleon: Fast-slow Neuro-symbolic Lane Topology Extraction
Lane topology extraction involves detecting lanes and traffic elements and determining their relationships, a key perception task for mapless autonomous driving. This task requires complex reasoning, such as determining …
Autonomous DrivingScene UnderstandingVisual PromptingTowards Ambiguity-Free Spatial Foundation Model: Rethinking and Decoupling Depth Ambiguity
Depth ambiguity is a fundamental challenge in spatial scene understanding, especially in transparent scenes where single-depth estimates fail to capture full 3D structure. Existing models, limited to deterministic predic…
Depth EstimationScene UnderstandingSpatial ReasoningVisual PromptingTowards Universal Text-driven CT Image Segmentation
Computed tomography (CT) is extensively used for accurate visualization and segmentation of organs and lesions. While deep learning models such as convolutional neural networks (CNNs) and vision transformers (ViTs) have …
Computed Tomography (CT)Contrastive LearningDiagnosticImage Segmentation+4The Role of Background Information in Reducing Object Hallucination in Vision-Language Models: Insights from Cutoff API Prompting
Vision-Language Models (VLMs) occasionally generate outputs that contradict input images, constraining their reliability in real-world applications. While visual prompting is reported to suppress hallucinations by augmen…
HallucinationObjectObject HallucinationVisual PromptingFrom PowerPoint UI Sketches to Web-Based Applications: Pattern-Driven Code Generation for GIS Dashboard Development Using Knowledge-Augmented LLMs, Context-Aware Visual Prompting, and the React Framework
Developing web-based GIS applications, commonly known as CyberGIS dashboards, for querying and visualizing GIS data in environmental research often demands repetitive and resource-intensive efforts. While Generative AI o…
Code GenerationRAGRetrieval-augmented GenerationVisual PromptingArticulate AnyMesh: Open-Vocabulary 3D Articulated Objects Modeling
3D articulated objects modeling has long been a challenging problem, since it requires to capture both accurate surface geometries and semantically meaningful and spatially precise structures, parts, and joints. Existing…
ObjectVisual PromptingPersonalization Toolkit: Training Free Personalization of Large Vision Language Models
Large Vision Language Models (LVLMs) have significant potential to deliver personalized assistance by adapting to individual users' unique needs and preferences. Personalization of LVLMs is an emerging area that involves…
RAGRetrievalRetrieval-augmented GenerationVisual PromptingLoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing visual prompting techniques often pad the p…
Inductive BiasVisual PromptingIP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual Prompting
The stories and characters that captivate us as we grow up shape unique fantasy worlds, with images serving as the primary medium for visually experiencing these realms. Personalizing generative models through fine-tunin…
Diffusion PersonalizationDiffusion Personalization Tuning FreeEfficient Diffusion PersonalizationImage Generation+4MedFocusCLIP : Improving few shot classification in medical datasets using pixel wise attention
With the popularity of foundational models, parameter efficient fine tuning has become the defacto approach to leverage pretrained models to perform downstream tasks. Taking inspiration from recent advances in large lang…
ClassificationFine-Grained Image Classificationimage-classificationImage Classification+4GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
In recent years, 2D Vision-Language Models (VLMs) have made significant strides in image-text understanding tasks. However, their performance in 3D spatial comprehension, which is critical for embodied intelligence, rema…
Scene Understandingtext annotationVisual PromptingQuery Efficient Black-Box Visual Prompting with Subspace Learning
Visual Prompt Learning (VPL) has emerged as a powerful strategy for harnessing the capabilities of large-scale pre-trained models (PTMs) to tackle specific downstream tasks. However, the opaque nature of PTMs in many…
Prompt LearningVisual PromptingVisual Prompting with Iterative Refinement for Design Critique Generation
Feedback is crucial for every design process, such as user interface (UI) design, and automating design critiques can significantly improve the efficiency of the design workflow. Although existing multimodal large langua…
AttributeVisual PromptingSelective Visual Prompting in Vision Mamba
Pre-trained Vision Mamba (Vim) models have demonstrated exceptional performance across various computer vision tasks in a computationally efficient manner, attributed to their unique design of selective state space model…
MambaState Space ModelsVisual PromptingTest-time Correction with Human Feedback: An Online 3D Detection System via Visual Prompting
This paper introduces Test-time Correction (TTC) system, a novel online 3D detection system designated for online correction of test-time errors via human feedback, to guarantee the safety of deployed autonomous driving …
Autonomous DrivingVisual PromptingInst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning
Large Multimodal Models (LMMs) have made significant breakthroughs with the advancement of instruction tuning. However, while existing models can understand images and videos at a holistic level, they still struggle with…
Multimodal Large Language ModelVideo Understandingvisual instruction followingVisual Prompting+1MLLM-Search: A Zero-Shot Approach to Finding People using Multimodal Large Language Models
Robotic search of people in human-centered environments, including healthcare settings, is challenging as autonomous robots need to locate people without complete or any prior knowledge of their schedules, plans or locat…
Person SearchVisual PromptingImproved GUI Grounding via Iterative Narrowing
Graphical User Interface (GUI) grounding plays a crucial role in enhancing the capabilities of Vision-Language Model (VLM) agents. While general VLMs, such as GPT-4V, demonstrate strong performance across various tasks, …
Language ModelingLanguage ModellingNatural Language Visual GroundingVisual PromptingPrompting the Unseen: Detecting Hidden Backdoors in Black-Box Models
Visual prompting (VP) is a new technique that adapts well-trained frozen models for source domain tasks to target domain tasks. This study examines VP's benefits for black-box model-level backdoor detection. The visual p…
Visual Prompting