paper-with-me

Papers Object Hallucination

“Object Hallucination” 태그가 달린 논문 71편 · 필터 해제

Revisit What You See: Disclose Language Prior in Vision Tokens for Efficient Guided Decoding of LVLMs

2025-06-11 · Beomsik Cho, Jaehyung Kim

Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across various multimodal tasks by integrating visual perception with language understanding. However, conventional decoding strategies of LVL…

HallucinationObject HallucinationText GenerationVisual Grounding

SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding

2025-06-10 · Woohyeon Park, Woojin Kim, Jaeik Kim, Jaeyoung Do

Despite significant advancements in Vision-Language Models (VLMs), the performance of existing VLMs remains hindered by object hallucination, a critical challenge to achieving accurate visual understanding. To address th…

HallucinationObject Hallucination

Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding

2025-06-08 · Tzu-wen Hsu, Ke-Han Lu, Cheng-Han Chiang, Hung-Yi Lee

Large Audio-Language Models (LALMs) can take audio and text as the inputs and answer questions about the audio. While prior LALMs have shown strong performance on standard benchmarks, there has been alarming evidence tha…

HallucinationObject Hallucination

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models

2025-05-26 · Jihoon Lee, Min Song

Despite significant advancements in Large Vision-Language Models, Object Hallucination (OH) remains a persistent challenge. Building upon prior studies on contrastive decoding that address this issue without requiring ad…

HallucinationObject HallucinationRetrieval

Visual Instruction Bottleneck Tuning

2025-05-20 · Changdae Oh, Jiatong Li, Shawn Im, Yixuan Li

Despite widespread adoption, multimodal large language models (MLLMs) suffer performance degradation when encountering unfamiliar queries under distribution shifts. Existing methods to improve MLLM generalization typical…

HallucinationObject HallucinationQuestion AnsweringRepresentation Learning

A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models

2025-05-04 · Liqiang Jing, Guiming Hardy Chen, Ehsan Aghazadeh, Xin Eric Wang 외

Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in multimodal tasks, but visual object hallucination remains a persistent issue. It refers to scenarios where models generate inaccurate visual obj…

AttributeHallucinationLanguage ModelingLanguage Modelling+3

Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models

2025-04-30 · Sangmin Woo, Kang Zhou, Yun Zhou, Shuai Wang 외

Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple object-based visual prompting -- overlaying visual cues (e.g., bounding b…

HallucinationObjectObject HallucinationPrompt Engineering+1

CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning

2025-03-25 · Hao Yu, Zhuokai Zhao, Shen Yan, Lukasz Korycki 외

The rapid advancement of large vision-language models (LVLMs) has driven significant progress in multimodal tasks, enabling models to interpret, reason, and generate outputs across both visual and textual domains. While …

HallucinationLanguage ModelingLanguage ModellingObject Hallucination+2

ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models

2025-03-17 · CVPR 2025 1 · Hao Yin, Guangzong Si, Zilei Wang

Contrastive decoding strategies are widely used to mitigate object hallucinations in multimodal large language models (MLLMs). By reducing over-reliance on language priors, these strategies ensure that generated content …

Computational EfficiencyHallucinationObject Hallucination

TruthPrInt: Mitigating LVLM Object Hallucination Via Latent Truthful-Guided Pre-Intervention

2025-03-13 · Jinhao Duan, Fei Kong, Hao Cheng, James Diffenderfer 외

Object Hallucination (OH) has been acknowledged as one of the major trustworthy challenges in Large Vision-Language Models (LVLMs). Recent advancements in Large Language Models (LLMs) indicate that internal states, such …

HallucinationObject HallucinationSpecificity

Seeing What's Not There: Spurious Correlation in Multimodal LLMs

2025-03-11 · Parsa Hosseini, Sumit Nawathe, Mazda Moayeri, Sriram Balasubramanian 외

Unimodal vision models are known to rely on spurious correlations, but it remains unclear to what extent Multimodal Large Language Models (MLLMs) exhibit similar biases despite language supervision. In this paper, we inv…

HallucinationObjectObject HallucinationObject Recognition

OmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal Inpainting

2025-03-11 · Yongsheng Yu, Ziyun Zeng, Haitian Zheng, Jiebo Luo

Diffusion-based generative models have revolutionized object-oriented image editing, yet their deployment in realistic object removal and insertion remains hampered by challenges such as the intricate interplay of physic…

HallucinationObjectObject Hallucination

EAZY: Eliminating Hallucinations in LVLMs by Zeroing out Hallucinatory Image Tokens

2025-03-10 · Liwei Che, Tony Qingze Liu, Jing Jia, Weiyi Qin 외

Despite their remarkable potential, Large Vision-Language Models (LVLMs) still face challenges with object hallucination, a problem where their generated outputs mistakenly incorporate objects that do not actually exist.…

HallucinationLanguage ModelingLanguage ModellingObject+1

Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow

2025-02-28 · Jiaqi Bai, Hongcheng Guo, Zhongyuan Peng, Jian Yang 외

Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffer from object hallucination, i.e., the generated image descriptions cont…

HallucinationObjectObject HallucinationSemantic Similarity+1

Vision-Encoders (Already) Know What They See: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore

2025-02-27 · Hongseok Oh, Wonseok Hwang

Recently, Large Vision-Language Models (LVLMs) show remarkable performance across various domains. However, these models suffer from object hallucination. This study revisits the previous claim that the primary cause of …

HallucinationObjectObject Hallucination

The Role of Background Information in Reducing Object Hallucination in Vision-Language Models: Insights from Cutoff API Prompting

2025-02-21 · Masayo Tomita, Katsuhiko Hayashi, Tomoyuki Kaneko

Vision-Language Models (VLMs) occasionally generate outputs that contradict input images, constraining their reliability in real-world applications. While visual prompting is reported to suppress hallucinations by augmen…

HallucinationObjectObject HallucinationVisual Prompting

CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base

2025-02-18 · Cong-Duy Nguyen, Xiaobao Wu, Duc Anh Vu, Shuai Zhao 외

Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, but they remain susceptible to hallucination, particularly object hallucination where non-existent objects or incorrect…

AttributeHallucinationMultimodal ReasoningObject Hallucination

Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration

2025-02-04 · Younan Zhu, Linwei Tao, Minjing Dong, Chang Xu

Large Vision-Language Models (LVLMs) exhibit impressive multimodal reasoning capabilities but remain highly susceptible to object hallucination, where models generate responses that are not factually aligned with the vis…

AttributeHallucinationMultimodal ReasoningObject+2

Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs

2025-01-31 · Kejia Zhang, Keda Tao, Jiasheng Tang, Huan Wang

Large vision-language models (LVMs) extend large language models (LLMs) with visual perception capabilities, enabling them to process and interpret visual information. A major challenge compromising their reliability is …

HallucinationObject Hallucination

Evaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities

2025-01-25 · Shounak Datta, Dhanasekar Sundararaman

Despite their impressive performance on multi-modal tasks, large vision-language models (LVLMs) tend to suffer from hallucinations. An important type is object hallucination, where LVLMs generate objects that are inconsi…

HallucinationObjectObject HallucinationObject Recognition
1–20 / 71 다음 →