paper-with-me

Papers

Dual Latent Memory for Visual Multi-agent System

2026-01-31 · Xinlei Yu, Chengming Xu, Zhangquan Chen, Bo Yin, Cheng Yang, Yongbo He, Yihao Hu, Jiangning Zhang, Cheng Tan, Xiaobin Hu, Shuicheng Yan arxiv

While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs. We attribute this failure to the information bottleneck inherent in text-centric communication, where converting perceptual and thinking trajectories into discrete natural language inevitably induces semantic loss. To this end, we propose \textbf{L}$\mathbf{^{2}}$\textbf{-VMAS}, a novel model-agnostic framework that enables inter-agent collaboration with dual latent memories. Furthermore, we decouple the perception and thinking while dynamically synthesizing dual latent memories. Additionally, we introduce an entropy-driven proactive triggering that replaces passive information transmission with efficient, on-demand memory access. Extensive experiments among backbones, sizes, and multi-agent structures demonstrate that our method effectively breaks the "scaling wall" with superb scalability, improving average accuracy by 2.7-5.4% while reducing token usage by 21.3-44.8%.

📄 PDF Abstract BibTeX arXiv:2602.00471

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection

2026-06-25 · Yujin Tang, Chenming Shang, Ruize Xu, Nikhil Singh arxiv

Agent benchmarks for measuring memory largely study textual cases, in which information is deliberately extracted from the environment, written down, and then later retrieved. In other words, they assess what agents elec…

LatentMem: Customizing Latent Memory for Multi-Agent Systems

2026-02-03 · Muxin Fu, Xiangyuan Xue, Yafu Li, Zefeng He 외 arxiv

Large language model (LLM)-powered multi-agent systems (MAS) demonstrate remarkable collective intelligence, wherein multi-agent memory serves as a pivotal mechanism for continual adaptation. However, existing multi-agen…

Personal Visual Memory from Explicit and Implicit Evidence

2026-05-27 · Viet Nguyen, Thao Nguyen, Vishal M. Patel, Yuheng Li arxiv

Long-term memory is increasingly important for personalized AI agents, yet existing benchmarks and methods remain largely text-centric. Even when images are included, the user-specific information needed for later questi…

Mem-W: Latent Memory-Native GUI Agents

2026-05-10 · Guibin Zhang, Yaohui Ling, Fanci Meng, Kun Wang 외 arxiv

GUI agents are beginning to operate the web, mobile, and desktop as interactive worlds, where successful control depends on carrying forward visual, procedural, and task-level evidence beyond the fleeting present screen.…

Agentic Learner with Grow-and-Refine Multimodal Semantic Memory

2025-11-26 · Weihao Bo, Shan Zhang, Yanpeng Sun, Jingjing Wu 외 arxiv

MLLMs exhibit strong reasoning on isolated queries, yet they operate de novo -- solving each problem independently and often repeating the same mistakes. Existing memory-augmented agents mainly store past trajectories fo…

Logical Reasoning