paper-with-me

홈 › Papers

Information-Theoretic Text Hallucination Reduction for Video-grounded Dialogue

2022-12-12 · Sunjae Yoon, Eunseop Yoon, Hee Suk Yoon, Junyeong Kim, Chang D. Yoo

Video-grounded Dialogue (VGD) aims to decode an answer sentence to a question regarding a given video and dialogue context. Despite the recent success of multi-modal reasoning to generate answer sentences, existing dialogue systems still suffer from a text hallucination problem, which denotes indiscriminate text-copying from input texts without an understanding of the question. This is due to learning spurious correlations from the fact that answer sentences in the dataset usually include the words of input texts, thus the VGD system excessively relies on copying words from input texts by hoping those words to overlap with ground-truth texts. Hence, we design Text Hallucination Mitigating (THAM) framework, which incorporates Text Hallucination Regularization (THR) loss derived from the proposed information-theoretic text hallucination measurement approach. Applying THAM with current dialogue systems validates the effectiveness on VGD benchmarks (i.e., AVSD@DSTC7 and AVSD@DSTC8) and shows enhanced interpretability.

📄 PDF Abstract BibTeX arXiv:2212.05765

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationSentence

Similar Papers 제목 키워드 기반

Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models

2024-01-18 · Li Sun, Liuan Wang, Jun Sun, Takayuki Okatani

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced the comprehension of multimedia content, bringing together diverse modalities such as text, images, and videos. However, a criti…

Hallucination

Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization

2026-02-03 · Hao Fang, Jinyu Li, Jiawei Kong, Tianqu Zhuang 외 arxiv

While multimodal reasoning models (MLRMs) have exhibited impressive capabilities, they remain prone to hallucinations, and effective solutions are still underexplored. In this paper, we experimentally analyze the halluci…

Multimodal Reasoning

On the Audio Hallucinations in Large Audio-Video Language Models

2024-01-18 · Taichi Nishimura, Shota Nakada, Masayoshi Kondo

Large audio-video language models can generate descriptions for both video and audio. However, they sometimes ignore audio content, producing audio descriptions solely reliant on visual information. This paper refers to …

HallucinationSentence

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

2025-08-29 · Hao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang 외 arxiv

Video multimodal large language models (Video-MLLMs) have achieved remarkable progress in video understanding. However, they remain vulnerable to hallucination-producing content inconsistent with or unrelated to video in…

SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse

2025-12-21 · Yiming Sun, Mi Zhang, Feifei Li, Geng Hong 외 arxiv

Despite Video Large Language Models having rapidly advanced in recent years, perceptual hallucinations pose a substantial safety risk, which severely restricts their real-world applicability. While several methods for ha…