paper-with-me

Papers

Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models

2026-03-16 · Amira Guesmi, Muhammad Shafique arxiv

Vision-language models (VLMs) have recently shown remarkable capabilities in visual understanding and generation, but remain vulnerable to adversarial manipulations of visual content. Prior object-hiding attacks primarily rely on suppressing or blocking region-specific representations, often creating semantic gaps that inadvertently induce hallucination, where models invent plausible but incorrect objects. In this work, we demonstrate that hallucination arises not from object absence per se, but from semantic discontinuity introduced by such suppression-based attacks. We propose a new class of \emph{background-consistent object concealment} attacks, which hide target objects by re-encoding their visual representations to be statistically and semantically consistent with surrounding background regions. Crucially, our approach preserves token structure and attention flow, avoiding representational voids that trigger hallucination. We present a pixel-level optimization framework that enforces background-consistent re-encoding across multiple transformer layers while preserving global scene semantics. Extensive experiments on state-of-the-art vision-language models show that our method effectively conceals target objects while preserving up to $86\%$ of non-target objects and reducing grounded hallucination by up to $3\times$ compared to attention-suppression-based attacks.

📄 PDF Abstract BibTeX arXiv:2603.15940

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations

2025-10-07 · Yike Wu, Yiwei Wang, Yujun Cai arxiv

While Large Vision-Language Models (LVLMs) achieve strong performance in multimodal tasks, hallucinations continue to hinder their reliability. Among the three categories of hallucinations, which include object, attribut…

Relational Reasoning

HulluEdit: Single-Pass Evidence-Consistent Subspace Editing for Mitigating Hallucinations in Large Vision-Language Models

2026-02-26 · Yangguang Lin, Quan Fang, Yufei Li, Jiachen Sun 외 arxiv

Object hallucination in Large Vision-Language Models (LVLMs) significantly hinders their reliable deployment. Existing methods struggle to balance efficiency and accuracy: they often require expensive reference models an…

Visual Grounding

THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models

2024-05-08 · CVPR 2024 1 · Prannay Kaul, Zhizhong Li, Hao Yang, Yonatan Dukler 외

Mitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which we term "Type I hallucinations". Instead…

AttributeData AugmentationFormHallucination+1

FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models

2023-11-02 · Liqiang Jing, Ruosen Li, Yunmo Chen, Xinya Du

We introduce FaithScore (Faithfulness to Atomic Image Facts Score), a reference-free and fine-grained evaluation metric that measures the faithfulness of the generated free-form answers from large vision-language models …

DescriptiveInstruction Following

Enhanced Spatially Interleaved Techniques for Multi-View Distributed Video Coding

2019-12-17 · Nantheera Anantrasirichai, Dimitris Agrafiotis

This paper presents a multi-view distributed video coding framework for independent camera encoding and centralized decoding. Spatio-temporal-view concealment methods are developed that exploit the interleaved nature of …

Diversity