Papers Object Hallucination
“Object Hallucination” 태그가 달린 논문 71편 · 필터 해제
Revisit What You See: Disclose Language Prior in Vision Tokens for Efficient Guided Decoding of LVLMs
Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across various multimodal tasks by integrating visual perception with language understanding. However, conventional decoding strategies of LVL…
HallucinationObject HallucinationText GenerationVisual GroundingSECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
Despite significant advancements in Vision-Language Models (VLMs), the performance of existing VLMs remains hindered by object hallucination, a critical challenge to achieving accurate visual understanding. To address th…
HallucinationObject HallucinationReducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
Large Audio-Language Models (LALMs) can take audio and text as the inputs and answer questions about the audio. While prior LALMs have shown strong performance on standard benchmarks, there has been alarming evidence tha…
HallucinationObject HallucinationRetrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
Despite significant advancements in Large Vision-Language Models, Object Hallucination (OH) remains a persistent challenge. Building upon prior studies on contrastive decoding that address this issue without requiring ad…
HallucinationObject HallucinationRetrievalVisual Instruction Bottleneck Tuning
Despite widespread adoption, multimodal large language models (MLLMs) suffer performance degradation when encountering unfamiliar queries under distribution shifts. Existing methods to improve MLLM generalization typical…
HallucinationObject HallucinationQuestion AnsweringRepresentation LearningA Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in multimodal tasks, but visual object hallucination remains a persistent issue. It refers to scenarios where models generate inaccurate visual obj…
AttributeHallucinationLanguage ModelingLanguage Modelling+3Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple object-based visual prompting -- overlaying visual cues (e.g., bounding b…
HallucinationObjectObject HallucinationPrompt Engineering+1CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
The rapid advancement of large vision-language models (LVLMs) has driven significant progress in multimodal tasks, enabling models to interpret, reason, and generate outputs across both visual and textual domains. While …
HallucinationLanguage ModelingLanguage ModellingObject Hallucination+2ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
Contrastive decoding strategies are widely used to mitigate object hallucinations in multimodal large language models (MLLMs). By reducing over-reliance on language priors, these strategies ensure that generated content …
Computational EfficiencyHallucinationObject HallucinationTruthPrInt: Mitigating LVLM Object Hallucination Via Latent Truthful-Guided Pre-Intervention
Object Hallucination (OH) has been acknowledged as one of the major trustworthy challenges in Large Vision-Language Models (LVLMs). Recent advancements in Large Language Models (LLMs) indicate that internal states, such …
HallucinationObject HallucinationSpecificitySeeing What's Not There: Spurious Correlation in Multimodal LLMs
Unimodal vision models are known to rely on spurious correlations, but it remains unclear to what extent Multimodal Large Language Models (MLLMs) exhibit similar biases despite language supervision. In this paper, we inv…
HallucinationObjectObject HallucinationObject RecognitionOmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal Inpainting
Diffusion-based generative models have revolutionized object-oriented image editing, yet their deployment in realistic object removal and insertion remains hampered by challenges such as the intricate interplay of physic…
HallucinationObjectObject HallucinationEAZY: Eliminating Hallucinations in LVLMs by Zeroing out Hallucinatory Image Tokens
Despite their remarkable potential, Large Vision-Language Models (LVLMs) still face challenges with object hallucination, a problem where their generated outputs mistakenly incorporate objects that do not actually exist.…
HallucinationLanguage ModelingLanguage ModellingObject+1Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow
Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffer from object hallucination, i.e., the generated image descriptions cont…
HallucinationObjectObject HallucinationSemantic Similarity+1Vision-Encoders (Already) Know What They See: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
Recently, Large Vision-Language Models (LVLMs) show remarkable performance across various domains. However, these models suffer from object hallucination. This study revisits the previous claim that the primary cause of …
HallucinationObjectObject HallucinationThe Role of Background Information in Reducing Object Hallucination in Vision-Language Models: Insights from Cutoff API Prompting
Vision-Language Models (VLMs) occasionally generate outputs that contradict input images, constraining their reliability in real-world applications. While visual prompting is reported to suppress hallucinations by augmen…
HallucinationObjectObject HallucinationVisual PromptingCutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, but they remain susceptible to hallucination, particularly object hallucination where non-existent objects or incorrect…
AttributeHallucinationMultimodal ReasoningObject HallucinationMitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
Large Vision-Language Models (LVLMs) exhibit impressive multimodal reasoning capabilities but remain highly susceptible to object hallucination, where models generate responses that are not factually aligned with the vis…
AttributeHallucinationMultimodal ReasoningObject+2Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs
Large vision-language models (LVMs) extend large language models (LLMs) with visual perception capabilities, enabling them to process and interpret visual information. A major challenge compromising their reliability is …
HallucinationObject HallucinationEvaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities
Despite their impressive performance on multi-modal tasks, large vision-language models (LVLMs) tend to suffer from hallucinations. An important type is object hallucination, where LVLMs generate objects that are inconsi…
HallucinationObjectObject HallucinationObject Recognition