paper-with-me

홈 › Papers

Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs

2025-02-21 · Gengyuan Zhang, Mingcong Ding, Tong Liu, Yao Zhang, Volker Tresp

Multimodal large language models (MLLMs) have demonstrated strong performance in understanding videos holistically, yet their ability to process streaming videos-videos are treated as a sequence of visual events-remains underexplored. Intuitively, leveraging past events as memory can enrich contextual and temporal understanding of the current event. In this paper, we show that leveraging memories as contexts helps MLLMs better understand video events. However, because such memories rely on predictions of preceding events, they may contain misinformation, leading to confabulation and degraded performance. To address this, we propose a confabulation-aware memory modification method that mitigates confabulated memory for memory-enhanced event understanding.

📄 PDF Abstract BibTeX arXiv:2502.15457

Code (0)

등록된 구현이 없습니다.

Tasks

Misinformation

Similar Papers 제목 키워드 기반

Honest Lying: Understanding Memory Confabulation in Reflexive Agents

2026-05-28 · Prakhar Dixit, Sadia Kamal, Tim Oates arxiv

Reflexion-style agents rely on self-generated reflections as memory, implicitly assuming that agents can accurately diagnose their own failures. We show that this assumption can fail systematically: across ALFWorld and H…

Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism

2026-03-31 · Tao Chen, Kun Zhang, Qiong Wu, Xiao Chen 외 arxiv

Long video understanding is a key challenge that plagues the advancement of \emph{Multimodal Large language Models} (MLLMs). In this paper, we study this problem from the perspective of visual memory mechanism, and propo…

Confabulation: The Surprising Value of Large Language Model Hallucinations

2024-06-06 · Peiqi Sui, Eamon Duede, Sophie Wu, Richard Jean So

This paper presents a systematic defense of large language model (LLM) hallucinations or 'confabulations' as a potential resource instead of a categorically negative pitfall. The standard view is that confabulations are …

HallucinationLanguage ModelingLanguage ModellingLarge Language Model+1

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

2025-10-13 · Guangzhi Sun, Yixuan Li, Xiaodong Wu, Yudong Yang 외 arxiv

Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhanced streaming audio-visual large language …

Critical Confabulation: Can LLMs Hallucinate for Social Good?

2025-11-11 · Peiqi Sui, Eamon Duede, Hoyt Long, Richard Jean So arxiv

LLMs hallucinate, yet some confabulations can have social affordances if carefully bounded. We propose critical confabulation (inspired by critical fabulation from literary and social theory), the use of LLM hallucinatio…