paper-with-me

Papers

Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models

2025-04-30 · Sangmin Woo, Kang Zhou, Yun Zhou, Shuai Wang, Sheng Guan, Haibo Ding, Lin Lee Cheong

Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple object-based visual prompting -- overlaying visual cues (e.g., bounding box, circle) on images -- can significantly mitigate such hallucination; however, different visual prompts (VPs) vary in effectiveness. To address this, we propose Black-Box Visual Prompt Engineering (BBVPE), a framework to identify optimal VPs that enhance LVLM responses without needing access to model internals. Our approach employs a pool of candidate VPs and trains a router model to dynamically select the most effective VP for a given input image. This black-box approach is model-agnostic, making it applicable to both open-source and proprietary LVLMs. Evaluations on benchmarks such as POPE and CHAIR demonstrate that BBVPE effectively reduces object hallucination.

📄 PDF Abstract BibTeX arXiv:2504.21559

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationObjectObject HallucinationPrompt EngineeringVisual Prompting

Similar Papers 제목 키워드 기반

Safer Prompts: Reducing IP Risk in Visual Generative AI

2025-05-06 · Lena Reissinger, Yuanyuan Li, Anna-Carolina Haensch, Neeraj Sarna

Visual Generative AI models have demonstrated remarkable capability in generating high-quality images from simple inputs like text prompts. However, because these models are trained on images from diverse sources, they r…

Image GenerationPrompt Engineering

Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation

2024-03-28 · Yutong He, Alexander Robey, Naoki Murata, Yiding Jiang 외

Prompt engineering is effective for controlling the output of text-to-image (T2I) generative models, but it is also laborious due to the need for manually crafted prompts. This challenge has spurred the development of al…

Image GenerationIn-Context LearningLanguage ModelingLanguage Modelling+4

Visual prompt engineering for video models

2026-07-28 · Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer, Neha Kalibhat 외 arxiv

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently bec…

Prompt EngineeringVisual ReasoningImage Editing

The Role of Background Information in Reducing Object Hallucination in Vision-Language Models: Insights from Cutoff API Prompting

2025-02-21 · Masayo Tomita, Katsuhiko Hayashi, Tomoyuki Kaneko

Vision-Language Models (VLMs) occasionally generate outputs that contradict input images, constraining their reliability in real-world applications. While visual prompting is reported to suppress hallucinations by augmen…

HallucinationObjectObject HallucinationVisual Prompting

Review of Large Vision Models and Visual Prompt Engineering

2023-07-03 · Jiaqi Wang, Zhengliang Liu, Lin Zhao, Zihao Wu 외

Visual prompt engineering is a fundamental technology in the field of visual and image Artificial General Intelligence, serving as a key component for achieving zero-shot capabilities. As the development of large vision …

Prompt Engineering