paper-with-me

홈 › Papers

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

2026-08-13 · Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao arxiv

Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Alaya-EVOKE (Evoke) addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation. Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as the session grows. Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons. Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence. A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term drift while preserving responsive conditioning. With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at $384\times 640$, each $1.5\,\mathrm{s}$ chunk is generated in $2.11\,\mathrm{s}$. As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.

📄 PDF Abstract BibTeX arXiv:2608.13546

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Endless Terminals: Scaling RL Environments for Terminal Agents

2026-01-23 · Kanishk Gandhi, Shivam Garg, Noah D. Goodman, Dimitris Papailiopoulos arxiv

Environments are the bottleneck for self-improving agents. Current terminal benchmarks were built for evaluation, not training; reinforcement learning requires a scalable pipeline, not just a dataset. We introduce Endles…

Reinforcement Learning

RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation

2023-11-02 · YuFei Wang, Zhou Xian, Feng Chen, Tsun-Hsuan Wang 외

We present RoboGen, a generative robotic agent that automatically learns diverse robotic skills at scale via generative simulation. RoboGen leverages the latest advancements in foundation and generative models. Instead o…

Motion Planning

A case study on English-Malayalam Machine Translation

2017-02-27 · Sreelekha. S, Pushpak Bhattacharyya

In this paper we present our work on a case study on Statistical Machine Translation (SMT) and Rule based machine translation (RBMT) for translation from English to Malayalam and Malayalam to English. One of the motivati…

Machine TranslationTranslation

How Useful is Self-Supervised Pretraining for Visual Tasks?

2020-03-31 · CVPR 2020 6 · Alejandro Newell, Jia Deng

Recent advances have spurred incredible progress in self-supervised pretraining for vision. We investigate what factors may play a role in the utility of these pretraining methods for practitioners. To do this, we evalua…

Linear evaluation

Stimuli reduce the dimensionality of cortical activity

2016-03-16

The activity of ensembles of simultaneously recorded neurons can be represented as a set of points in the space of firing rates. Even though the dimension of this space is equal to the ensemble size, neural activity can …