paper-with-me

홈 › Papers

RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models

2024-05-28 · Sangmin Woo, Jaehyuk Jang, Donguk Kim, Yubin Choi, Changick Kim

Recent advancements in Large Vision Language Models (LVLMs) have revolutionized how machines understand and generate textual responses based on visual inputs, yet they often produce "hallucinatory" outputs that misinterpret visual information, posing challenges in reliability and trustworthiness. We propose RITUAL, a simple decoding method that reduces hallucinations by leveraging randomly transformed images as complementary inputs during decoding, adjusting the output probability distribution without additional training or external models. Our key insight is that random transformations expose the model to diverse visual perspectives, enabling it to correct misinterpretations that lead to hallucinations. Specifically, when a model hallucinates based on the original image, the transformed images -- altered in aspects such as orientation, scale, or color -- provide alternative viewpoints that help recalibrate the model's predictions. By integrating the probability distributions from both the original and transformed images, RITUAL effectively reduces hallucinations. To further improve reliability and address potential instability from arbitrary transformations, we introduce RITUAL+, an extension that selects image transformations based on self-feedback from the LVLM. Instead of applying transformations randomly, RITUAL+ uses the LVLM to evaluate and choose transformations that are most beneficial for reducing hallucinations in a given context. This self-adaptive approach mitigates the potential negative impact of certain transformations on specific tasks, ensuring more consistent performance across different scenarios. Experiments demonstrate that RITUAL and RITUAL+ significantly reduce hallucinations across several object hallucination benchmarks.

📄 PDF Abstract BibTeX arXiv:2405.17821

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationMMEObject Hallucination

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Data-free Universal Adversarial Perturbation with Pseudo-semantic Prior

2025-02-28 · CVPR 2025 1 · Chanhui Lee, Yeonghwan Song, Jeany Son

Data-free Universal Adversarial Perturbation (UAP) is an image-agnostic adversarial attack that deceives deep neural networks using a single perturbation generated solely from random noise without relying on data priors.…

Adversarial Attack

Experimental Tests of Spirituality

2018-06-04 · Abraham Loeb

We currently harness technologies that could shed new light on old philosophical questions, such as whether our mind entails anything beyond our body or whether our moral values reflect universal truth.

Measuring Spiritual Values and Bias of Large Language Models

2024-10-15 · Songyuan Liu, Ziyang Zhang, Runze Yan, Wei Wu 외

Large language models (LLMs) have become integral tool for users from various backgrounds. LLMs, trained on vast corpora, reflect the linguistic and cultural nuances embedded in their pre-training data. However, the valu…

Fairness

Enhancing Transferability of Targeted Adversarial Examples: A Self-Universal Perspective

2024-07-22 · Bowen Peng, Li Liu, Tianpeng Liu, Zhen Liu 외

Transfer-based targeted adversarial attacks against black-box deep neural networks (DNNs) have been proven to be significantly more challenging than untargeted ones. The impressive transferability of current SOTA, the ge…

Matrix optimization on universal unitary photonic devices

2018-08-02 · Sunil Pai, Ben Bartlett, Olav Solgaard, David A. B. Miller

Universal unitary photonic devices can apply arbitrary unitary transformations to a vector of input modes and provide a promising hardware platform for fast and energy-efficient machine learning using light. We simulate …