paper-with-me

홈 › Papers

Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation

2026-06-03 · Yuxuan Bian, Zeyue Xue, Songchun Zhang, Shiyi Zhang, Weiyang Jin, Yaowei Li, Junhao Zhuang, Haoran Li, Jie Huang, Haoyang Huang, Nan Duan, Qiang Xu arxiv

We present Echo Infinity, an autoregressive (AR) framework towards real-time infinite video generation that employs a learnable evolving memory to dynamically filter, abstract, and compress any-length history at constant cost. Existing methods mainly curate memory with predefined KV-cache schedules, fixed-ratio heuristic compression, or inference-time RoPE adaptation. These designs inevitably lose historical information and amplify compounding errors due to their limited cache window and ignorance of autoregressive generation noise. Inspired by human memory consolidation, Echo-Infinity replaces handcrafted memory curation with learnable Memory Query, which are updated by attention and a gating mechanism when past frames are evicted from the local window. The queries are optimized end-to-end with the video diffusion transformers (DiTs), forming an evolving memory that supports arbitrary compression ratios with constant computation independent of video length. They also act as a generalizable generation prior, improving quality even when only the optimized initial state is used. We further introduce Unified Relative RoPE Recipe, which anchors the sink frames to start from id 0 and lets the newest frame id grow at most to the DiTs' pretrained maximum temporal RoPE id throughout training and inference, freeing the model from the finite RoPE constraint and closing the train-test RoPE extrapolation gap. In long and short video generation, Echo-Infinity achieves state-of-the-art performance, and, to our knowledge, demonstrates promising 24-hour (>1.3 M frames) real-time rollouts for the first time, suggesting a practical path toward infinite video generation.

📄 PDF Abstract BibTeX arXiv:2606.04527

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

InfinityDrive: Breaking Time Limits in Driving World Models

2024-12-02 · Xi Guo, Chenjing Ding, Haoxuan Dou, Xin Zhang 외

Autonomous driving systems struggle with complex scenarios due to limited access to diverse, extensive, and out-of-distribution driving data which are critical for safe navigation. World models offer a promising solution…

Autonomous DrivingDiversityVideo Generation

Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System

2024-03-11 · Hongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park 외

The recent huge advance of Large Language Models (LLMs) is mainly driven by the increase in the number of parameters. This has led to substantial memory capacity requirements, necessitating the use of dozens of GPUs just…

GPULanguage ModelingLanguage ModellingLarge Language Model

Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory

2026-03-16 · Ce Zhang, Jinxi He, Junyi He, Katia Sycara 외 arxiv

Multi-modal Large Language Models (MLLMs) have achieved remarkable performance across a wide range of visual reasoning tasks, yet their vulnerability to safety risks remains a pressing concern. While prior research prima…

Visual Reasoning

Nonlinear Residual Echo Suppression using a Recurrent Neural Network

2020-10-25 · Interspeech 2020 10 · Lukas Pfeifenberger, Franz Pernkopf

The acoustic front-end of hands-free communication de-vices introduces a variety of distortions to the linear echo pathbetween the loudspeaker and the microphone. While the ampli-fiers may introduce a memory-less n…

Acoustic echo cancellation

Echo state networks are universal

2018-06-03 · Lyudmila Grigoryeva, Juan-Pablo Ortega

This paper shows that echo state networks are universal uniform approximants in the context of discrete-time fading memory filters with uniformly bounded inputs defined on negative infinite times. This result guarantees …

valid