paper-with-me

홈 › Papers

BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models

2024-07-18 · Moon Ye-Bin, Nam Hyeon-Woo, Wonseok Choi, Tae-Hyun Oh

Vision language models (VLMs) perceive the world through a combination of a visual encoder and a large language model (LLM). The visual encoder, pre-trained on large-scale vision-text datasets, provides zero-shot generalization to visual data, and the LLM endows its high reasoning ability to VLMs. It leads VLMs to achieve high performance on wide benchmarks without fine-tuning, exhibiting zero or few-shot capability. However, recent studies show that VLMs are vulnerable to hallucination. This undesirable behavior degrades reliability and credibility, thereby making users unable to fully trust the output from VLMs. To enhance trustworthiness and better tackle the hallucination of VLMs, we curate a new evaluation dataset, called the BEfore-AFter hallucination dataset (BEAF), and introduce new metrics: True Understanding (TU), IGnorance (IG), StuBbornness (SB), and InDecision (ID). Unlike prior works that focus only on constructing questions and answers, the key idea of our benchmark is to manipulate visual scene information by image editing models and to design the metrics based on scene changes. This allows us to clearly assess whether VLMs correctly understand a given scene by observing the ability to perceive changes. We also visualize image-wise object relationship by virtue of our two-axis view: vision and text. Upon evaluating VLMs with our dataset, we observed that our metrics reveal different aspects of VLM hallucination that have not been reported before. Project page: \url{https://beafbench.github.io/}

📄 PDF Abstract BibTeX arXiv:2407.13442

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationLanguage ModellingLarge Language ModelZero-shot Generalization

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Analyzing Privacy Loss in Updates of Natural Language Models

2019-09-25 · Shruti Tople, Marc Brockschmidt, Boris Köpf, Olga Ohrimenko 외

To continuously improve quality and reflect changes in data, machine learning-based services have to regularly re-train and update their core models. In the setting of language models, we show that a comparative analysis…

Unobtrusive Monitoring of Physical Weakness: A Simulated Approach

2024-06-14 · Chen Long-fei, Muhammad Ahmed Raza, Craig Innes, Subramanian Ramamoorthy 외

Aging and chronic conditions affect older adults' daily lives, making early detection of developing health issues crucial. Weakness, common in many conditions, alters physical movements and daily activities subtly. Howev…

Fully adaptive algorithm for pure exploration in linear bandits

2017-10-16 · Liyuan Xu, Junya Honda, Masashi Sugiyama

We propose the first fully-adaptive algorithm for pure exploration in linear bandits---the task to find the arm with the largest expected reward, which depends on an unknown parameter linearly. While existing methods par…

Optimal Pre-Processing to Achieve Fairness and Its Relationship with Total Variation Barycenter

2021-01-18 · Farhad Farokhi

We use disparate impact, i.e., the extent that the probability of observing an output depends on protected attributes such as race and gender, to measure fairness. We prove that disparate impact is upper bounded by the t…

Fairness

Mean Field Game of Controls and An Application To Trade Crowding

2017-09-21

In this paper we formulate the now classical problem of optimal liquidation (or optimal trading) inside a Mean Field Game (MFG). This is a noticeable change since usually mathematical frameworks focus on one large trader…