paper-with-me

Papers

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models

2024-12-06 · Zehao Wang, Xinpeng Liu, Xiaoqian Wu, Yudonglin Zhang, Zhou Fang, Yifan Fang, Junfu Pu, Cewu Lu, Yong-Lu Li

Multimodal Large Language Models (MLLMs) have garnered significant attention recently and demonstrate outstanding capabilities in various tasks such as OCR, VQA, captioning, $\textit{etc}$. However, hallucination remains a persistent issue. While numerous methods have been proposed to mitigate hallucinations, achieving notable improvements, these methods primarily focus on mitigating hallucinations about $\textbf{object/noun-related}$ concepts. Verb concepts, crucial for understanding human actions, have been largely overlooked. In this paper, to the best of our knowledge, we are the $\textbf{first}$ to investigate the $\textbf{verb hallucination}$ phenomenon of MLLMs from various perspectives. Our findings reveal that most state-of-the-art MLLMs suffer from severe verb hallucination. To assess the effectiveness of existing mitigation methods for object concept hallucination on verb hallucination, we evaluated these methods and found that they do not effectively address verb hallucination. To address this issue, we propose a novel rich verb knowledge-based tuning method to mitigate verb hallucination. The experiment results demonstrate that our method significantly reduces hallucinations related to verbs. $\textit{Our code and data will be made publicly available}$.

📄 PDF Abstract BibTeX arXiv:2412.04939

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationOptical Character Recognition (OCR)Visual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses

2026-06-01 · Sugyeong Eo, Heuiseok Lim arxiv

Although large language models (LLMs) have shown considerable progress in pragmatic language understanding, prior research has focused mainly on their comprehension of verbal behavior. Nonetheless, non-verbal behavior re…

Representing Verbs as Argument Concepts

2018-03-02 · Yu Gong, Kaiqi Zhao, Kenny Q. Zhu

Verbs play an important role in the understanding of natural language text. This paper studies the problem of abstracting the subject and object arguments of a verb into a set of noun concepts, known as the "argument con…

Object

Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model

2025-01-21 · Minghan Wang, Viet-Thanh Pham, Farhad Moghimifar, Thuy-Trang Vu

Despite achieving remarkable performance, machine translation (MT) research remains underexplored in terms of translating cultural elements in languages, such as idioms, proverbs, and colloquial expressions. This paper i…

Language ModelingLanguage ModellingLarge Language ModelMachine Translation+2

Verb Pattern: A Probabilistic Semantic Representation on Verbs

2017-10-20 · Wanyun Cui, Xiyou Zhou, Hangyu Lin, Yanghua Xiao 외

Verbs are important in semantic understanding of natural language. Traditional verb representations, such as FrameNet, PropBank, VerbNet, focus on verbs' roles. These roles are too coarse to represent verbs' semantics. I…

Specificity

A multimodal interpreter for 3D visualization and animation of verbal concepts

2014-05-01 · LREC 2014 5 · Coline Claude-Lachenaud, {\'E}ric Charton, Beno{\^\i}t Ozell, Michel Gagnon

We present an algorithm intended to visually represent the sense of verb related to an object described in a text sequence, as a movement in 3D space. We describe a specific semantic analyzer, based on a standard verbal …

Text Generation