paper-with-me

홈 › Papers

Understanding and Evaluating Hallucinations in 3D Visual Language Models

2025-02-18 · Ruiying Peng, Kaiyuan Li, Weichen Zhang, Chen Gao, Xinlei Chen, Yong Li

Recently, 3D-LLMs, which combine point-cloud encoders with large models, have been proposed to tackle complex tasks in embodied intelligence and scene understanding. In addition to showing promising results on 3D tasks, we found that they are significantly affected by hallucinations. For instance, they may generate objects that do not exist in the scene or produce incorrect relationships between objects. To investigate this issue, this work presents the first systematic study of hallucinations in 3D-LLMs. We begin by quickly evaluating hallucinations in several representative 3D-LLMs and reveal that they are all significantly affected by hallucinations. We then define hallucinations in 3D scenes and, through a detailed analysis of datasets, uncover the underlying causes of these hallucinations. We find three main causes: (1) Uneven frequency distribution of objects in the dataset. (2) Strong correlations between objects. (3) Limited diversity in object attributes. Additionally, we propose new evaluation metrics for hallucinations, including Random Point Cloud Pair and Opposite Question Evaluations, to assess whether the model generates responses based on visual information and aligns it with the text's meaning.

📄 PDF Abstract BibTeX arXiv:2502.15888

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityScene Understanding

Similar Papers 제목 키워드 기반

Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models

2024-06-24 · Mingrui Wu, Jiayi Ji, Oucheng Huang, Jiale Li 외

The issue of hallucinations is a prevalent concern in existing Large Vision-Language Models (LVLMs). Previous efforts have primarily focused on investigating object hallucinations, which can be easily alleviated by intro…

Common Sense ReasoningHallucinationObject

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

2024-12-04 · CVPR 2025 1 · Chaoyu Li, Eun Woo Im, Pooyan Fazli

Multimodal large language models (MLLMs) have recently shown significant advancements in video understanding, excelling in content reasoning and instruction-following tasks. However, hallucination, where models generate …

HallucinationInstruction FollowingSemantic SimilaritySemantic Textual Similarity+1

AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

2024-10-23 · Kim Sung-Bin, Oh Hyun-Bin, JungMok Lee, Arda Senocak 외

Following the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying …

Hallucination

VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

2024-06-24 · Yuxuan Wang, Yueqian Wang, Dongyan Zhao, Cihang Xie 외

Recent advancements in Multimodal Large Language Models (MLLMs) have extended their capabilities to video understanding. Yet, these models are often plagued by "hallucinations", where irrelevant or nonsensical content is…

HallucinationVideo Understanding

Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models

2025-06-25 · Zhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen 외

Recent advancements in multimodal large language models have enhanced document understanding by integrating textual and visual information. However, existing models exhibit incompleteness within their paradigm in real-wo…

document understandingHallucinationOptical Character Recognition (OCR)