paper-with-me

홈 › Papers

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

2026-05-14 · Minghao Guo, Qingyue Jiao, Zeru Shi, Yihao Quan, Boxuan Zhang, Danrui Li, Liwei Che, Wujiang Xu, Shilong Liu, Zirui Liu, Mubbasir Kapadia, Vladimir Pavlovic, Jiang Liu, Mengdi Wang, Yiyu Shi, Dimitris N. Metaxas, Ruixiang Tang arxiv

Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later reasoning. In prior work, many visually grounded questions can be answered using only captions or textual traces, allowing answers to be inferred without preserving the fine-grained visual evidence. Meanwhile, harder cases that require reasoning over changing visual states are largely absent. Therefore, we introduce MemEye, a framework that evaluates memory capabilities from two dimensions: one measures the granularity of decisive visual evidence (from scene-level to pixel-level evidence), and the other measures how retrieved evidence must be used (from single evidence to evolutionary synthesis). Under this framework, we construct a new benchmark across 8 life-scenario tasks, with ablation-driven validation gates for assessing answerability, shortcut resistance, visual necessity, and reasoning structure. By evaluating 13 memory methods across 4 VLM backbones, we show that current architectures still struggle to preserve fine-grained visual details and reason about state changes over time. Our findings show that long-term multimodal memory depends on evidence routing, temporal tracking, and detail extraction.

📄 PDF Abstract BibTeX arXiv:2605.15128

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation

2025-10-28 · Alexander Martin, William Walden, Reno Kriz, Dengjia Zhang 외 arxiv

We introduce MiRAGE, an evaluation framework for retrieval-augmented generation (RAG) from multimodal sources. As audiovisual media becomes a prevalent source of information online, it is essential for RAG systems to int…

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

2025-05-29 · CVPR 2025 1 · Yunze Man, De-An Huang, Guilin Liu, Shiwei Sheng 외

Recent advances in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks, yet they often struggle with vision-centric scenarios where precise visual focus is needed f…

Multimodal Reasoning

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

2026-04-13 · Yuqian Yuan, Wenqiao Zhang, Juekai Lin, Yu Zhong 외 arxiv

Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasonin…

Scene UnderstandingSpatial Reasoning

Beyond Retrieval: Analytic Memory for Multimodal Agents

2026-07-31 · Zhoujin Tian, Hao Zhang, Yao Tian, Cheng Chen 외 arxiv

Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizi…

Exploring Audio Hallucination in Egocentric Video Understanding

2026-04-26 · Ashish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja 외 arxiv

Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly when visual information is unstable or occluded due to continuous camera …