paper-with-me

홈 › Papers

RWKV-7 "Goose" with Expressive Dynamic State Evolution

2025-03-18 · Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Xingjian Du, Haowen Hou, Jiaju Lin, Jiaxing Liu, Janna Lu, William Merrill, Guangyu Song, Kaifeng Tan, Saiteja Utpala, Nathan Wilce, Johan S. Wind, Tianyi Wu, Daniel Wuttke, Christian Zhou-Zheng

We present RWKV-7 "Goose", a new sequence modeling architecture with constant memory usage and constant inference time per token. Despite being trained on dramatically fewer tokens than other top models, our 2.9 billion parameter language model achieves a new 3B SoTA on multilingual tasks and matches the current 3B SoTA on English language downstream performance. RWKV-7 introduces a newly generalized formulation of the delta rule with vector-valued gating and in-context learning rates, as well as a relaxed value replacement rule. We show that RWKV-7 can perform state tracking and recognize all regular languages, while retaining parallelizability of training. This exceeds the capabilities of Transformers under standard complexity conjectures, which are limited to $\mathsf{TC}^0$. To demonstrate RWKV-7's language modeling capability, we also present an extended open source 3.1 trillion token multilingual corpus, and train four RWKV-7 models ranging from 0.19 billion to 2.9 billion parameters on this dataset. To foster openness, reproduction, and adoption, we release our models and dataset component listing at https://huggingface.co/RWKV, and our training and inference code at https://github.com/RWKV/RWKV-LM all under the Apache 2.0 License.

📄 PDF Abstract BibTeX arXiv:2503.14456

Code (3)

fla-org/flash-linear-attention 공식 구현 pytorch
rwkv/rwkv-lm 공식 구현 pytorch
blinkdl/chatrwkv pytorch

Tasks

In-Context LearningLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

BlackGoose Rimer: Harnessing RWKV-7 as a Simple yet Superior Replacement for Transformers in Large-Scale Time Series Modeling

2025-03-08 · Li weile, Liu Xiao

Time series models face significant challenges in scaling to handle large and complex datasets, akin to the scaling achieved by large language models (LLMs). The unique characteristics of time series data and the computa…

Meta-LearningTime Series

ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer

2025-01-26 · Lin Yueyu, Li Zhiyuan, Peter Yue, Liu Xiao

As is known, hybrid quadratic and subquadratic attention models in multi-head architectures have surpassed both Transformer and Linear RNN models , with these works primarily focusing on reducing KV complexity and improv…

Language ModelingLanguage ModellingTransfer Learning

Cross-attention for State-based model RWKV-7

2025-04-19 · Liu Xiao, Li Zhiyuan, Lin Yueyu

We introduce CrossWKV, a novel cross-attention mechanism for the state-based RWKV-7 model, designed to enhance the expressive power of text-to-image generation. Leveraging RWKV-7's linear-complexity Weighted Key-Value (W…

cross-modal alignmentImage GenerationmodelText to Image Generation+1

PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning

2024-05-24 · Qingdong He, Jiangning Zhang, Jinlong Peng, Haoyang He 외

Transformers have revolutionized the point cloud learning task, but the quadratic complexity hinders its extension to long sequence and makes a burden on limited computational resources. The recent advent of RWKV, a fres…

Mamba

Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

2024-04-08 · Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak 외

We present Eagle (RWKV-5) and Finch (RWKV-6), sequence models improving upon the RWKV (RWKV-4) architecture. Our architectural design advancements include multi-headed matrix-valued states and a dynamic recurrence mechan…