paper-with-me

홈 › Papers

Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering

2025-05-19 · JianFeng Cai, Wengang Zhou, Zongmeng Zhang, Jiale Hong, Nianji Zhan, Houqiang Li

Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding.However, hallucination, where the model generates plausible yet incorrect outputs, persists as a significant and under-addressed challenge in the video domain. Among existing solutions, activation engineering has proven successful in mitigating hallucinations in LLMs and ImageLLMs, yet its applicability to VideoLLMs remains largely unexplored. In this work, we are the first to systematically investigate the effectiveness and underlying mechanisms of activation engineering for mitigating hallucinations in VideoLLMs. We initially conduct an investigation of the key factors affecting the performance of activation engineering and find that a model's sensitivity to hallucination depends on $\textbf{temporal variation}$ rather than task type. Moreover, selecting appropriate internal modules and dataset for activation engineering is critical for reducing hallucination. Guided by these findings, we propose a temporal-aware activation engineering framework for VideoLLMs, which adaptively identifies and manipulates hallucination-sensitive modules based on the temporal variation characteristic, substantially mitigating hallucinations without additional LLM fine-tuning. Experiments across multiple models and benchmarks demonstrate that our method markedly reduces hallucination in VideoLLMs, thereby validating the robustness of our findings.

📄 PDF Abstract BibTeX arXiv:2505.12826

Code (0)

등록된 구현이 없습니다.

Tasks

Hallucination

Similar Papers 제목 키워드 기반

SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding

2025-12-04 · Chang-Hsun Wu, Kai-Po Chang, Yu-Yang Sheng, Hung-Kai Chung 외 arxiv

Video Large Language Models (VideoLLMs) have shown remarkable progress in video understanding. However, these models still struggle to effectively perceive and exploit rich temporal information in videos when responding …

EventHallusion: Diagnosing Event Hallucinations in Video LLMs

2024-09-25 · Jiacheng Zhang, Yang Jiao, Shaoxiang Chen, Na Zhao 외

Recently, Multimodal Large Language Models (MLLMs) have made significant progress in the video comprehension field. Despite remarkable content reasoning and instruction following capabilities they demonstrated, the hallu…

HallucinationInstruction Following

VERHallu: Evaluating and Mitigating Event Relation Hallucination in Video Large Language Models

2026-01-15 · Zefan Zhang, Kehua Zhu, Shijie Jiang, Hongyuan Lu 외 arxiv

Video Large Language Models (VideoLLMs) exhibit various types of hallucinations. Existing research has primarily focused on hallucinations involving the presence of events, objects, and scenes in videos, while largely ne…

Relation ClassificationQuestion Answering

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models

2026-07-01 · Jiale Li, Sihan Chen, Mengyuan Liu arxiv

Video Large Language Models (VideoLLMs) have shown strong progress in video understanding, yet they still suffer from hallucinations that are inconsistent with visual evidence. Existing benchmarks mainly focus on object …

Action Recognition

Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models

2026-05-26 · Oscar Chew, Serhii Honcharenko, Qian-Hui Chen, Patricia Lu 외 arxiv

A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In this work, we introduce Distracti…