paper-with-me

Papers

Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations

2026-05-27 · Kevin Y. Li, Asher Trockman, Ananda Theertha Suresh, Ziteng Sun arxiv

Softmax attention is the cornerstone of modern large language models, but its memory scales linearly and compute quadratically with sequence length. Linear recurrent models, such as linear attention and state space models, have become widely studied as alternatives to attention due to their linear compute and constant memory. While these sub-quadratic token mixing methods, or mixers, achieve promising efficiency gains and competitive results on a wide range of benchmarks, current linear recurrent models still lag behind on tasks that require long-context retrieval or in-context learning. A growing body of work studies hybrid architectures that attempt to mitigate these trade-offs by statically interleaving or merging attention and recurrent blocks. In this work, we explore a new axis of developing hybrid models: across the token sequence. We propose Oryx, a hybrid model that can, throughout a sequence, flexibly switch between different mixers, for example quadratic attention for rich context utilization and linear recurrences for efficient generation. Oryx ties at least 90% of its parameters across mixers, enabling attention and recurrent modes to operate over shared internal representations. We validate our design with Mamba-2 and Gated DeltaNet variants, up to 1.4B models. Under fixed token budgets and a mixed-training strategy, Oryx achieves comparable or better performance than its single-mixer baselines. At the 1.4B scale, all instances of Oryx outperform their respective baselines by at least 0.7 percentage points on averaged language modeling tasks. On retrieval tasks, Oryx achieves performance comparable to the Transformer baseline even when processing only a tiny fraction (<10%) of the tokens in attention mode. These results suggest that attention and linear recurrent models can share internal representations, and motivate sequence-axis hybridization as a promising direction.

📄 PDF Abstract BibTeX arXiv:2605.28769

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

xLSTM-Mixer: Multivariate Time Series Forecasting by Mixing via Scalar Memories

2024-10-22 · Maurice Kraus, Felix Divo, Devendra Singh Dhami, Kristian Kersting

Time series data is prevalent across numerous fields, necessitating the development of robust and accurate forecasting models. Capturing patterns both within and between temporal and multivariate components is crucial fo…

Multivariate Time Series ForecastingTemporal SequencesTime SeriesTime Series Forecasting

MTS-UNMixers: Multivariate Time Series Forecasting via Channel-Time Dual Unmixing

2024-11-26 · Xuanbing Zhu, Dunbin Shen, Zhongwen Rao, Huiyi Ma 외

Multivariate time series data provide a robust framework for future predictions by leveraging information across multiple dimensions, ensuring broad applicability in practical scenarios. However, their high dimensionalit…

MambaMultivariate Time Series ForecastingTime SeriesTime Series Forecasting

Field-Aware RankMixer with Dual-Stream Bilinear Fusion for the Tencent UNI-REC Challenge

2026-07-17 · Yufeng Zhang, Zhengqi Xu, Jiajun Cui arxiv

This paper presents our solution to the KDD Cup 2026 Tencent UNIREC Challenge. The task requires joint modeling of multi-domain user behavior sequences and non-sequential multi-field features for target-ad pCVR predictio…

SDMixer: Sparse Dual-Mixer for Time Series Forecasting

2026-02-27 · Xiang Ao arxiv

Multivariate time series forecasting is widely applied in fields such as transportation, energy, and finance. However, the data commonly suffers from issues of multi-scale characteristics, weak correlations, and noise in…

Multivariate Time Series Forecasting

UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation

2026-08-17 · Rongcheng Lin, Yan Sun, Jamey Zhang, Guanglei Xiong 외 arxiv

Industrial recommenders rely on two model families that have evolved largely independently: feature-interaction models over multi-field user/item features, and sequential models over user-behavior histories. Production s…

Collaborative Filtering