paper-with-me

Papers

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs

2026-05-01 · Siyuan Huang, Xiaoye Qu, Yafu Li, Tong Zhu, Zefeng He, Muxin Fu, Daizong Liu, Wei-Long Zheng, Yu Cheng arxiv

While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a "Visual Signal Dilution" phenomenon, where the accumulation of textual history expands the attention partition function, causing visual attention to decay inversely with generated sequence length. To counteract this, we propose Persistent Visual Memory (PVM), a lightweight learnable module designed to strengthen sustained, on-demand access to visual evidence. Integrated as a parallel branch alongside the Feed-Forward Network (FFN) in LVLMs, PVM establishes a distance-agnostic retrieval pathway that directly provides visual embeddings for enhanced visual perception, thereby structurally mitigating the signal suppression inherent to deep generation. Extensive experiments on Qwen3-VL models demonstrate that PVM brings notable improvements with negligible parameter overhead, delivering consistent average accuracy gains across both 4B and 8B scales, particularly in complex reasoning tasks that demand persistent visual perception. Furthermore, in-depth analysis reveals that PVM shows improved robustness in longer generations and accelerates internal prediction convergence.

📄 PDF Abstract BibTeX arXiv:2605.00814

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusion

2026-06-02 · Matiur Rahman Minar, Seunghun Oh, GangHyeon Jeong, Unsang Park arxiv

Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene layouts drift, while mechanisms that improve spatial stability tend to suppress motion, causing natural…

Video Generation

Networking-Aware Energy Efficiency in Agentic AI Inference: A Survey

2026-04-09 · Xiaojing Chen, Haiqi Yu, Wei Ni, Dusit Niyato 외 arxiv

The rapid emergence of Large Language Models (LLMs) has catalyzed Agentic artificial intelligence (AI), autonomous systems integrating perception, reasoning, and action into closed-loop pipelines for continuous adaptatio…

Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping

2026-06-03 · Peilin Tao, Chong Cheng, Yuansen Du, Caiwei Song 외 arxiv

Long-horizon online visual mapping is a core capability for robot perception, requiring continuous camera-motion and scene-geometry estimation from visual streams under bounded memory and computation. Recent feed-forward…

3D Reconstruction

WorldKV: Efficient World Memory with World Retrieval and Compression

2026-05-21 · Jung Yi, Minjae Kim, Paul Hyunbin Cho, Wooseok Jang 외 arxiv

Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpoint yields consistent content, remains a…

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

2026-07-02 · Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang 외 arxiv

We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynami…

Video Generation