paper-with-me

홈 › Papers

An Integrable Token Mixing Layer from the Generalized Yang Baxter Equation

2026-06-13 · Snigdha Chandan Khilar arxiv

The YB Mixer is a sequence token mixing layer derived from free fermion and generalized Yang Baxter structures. It applies a core principle from integrable systems where a local algebraic constraint guarantees global computational stability. By using the Ising exchange algebra the mixer creates a free fermionic structure that acts as an exactly norm preserving orthogonal map. This algebra also produces commuting transfer matrices which allow inference to be order free and adaptable to any variable budget. To ensure the model can generalize to longer sequence lengths it uses a spectral circulant generator. This generator maintains the crucial orthogonal and commuting properties of the system. The result is a highly stable and mathematically grounded architecture for sequence processing.

📄 PDF Abstract BibTeX arXiv:2606.15085

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The R-mAtrIx Net

2023-04-14 · Shailesh Lal, Suvajit Majumder, Evgeny Sobko

We provide a novel Neural Network architecture that can: i) output R-matrix for a given quantum integrable spin chain, ii) search for an integrable Hamiltonian and the corresponding R-matrix under assumptions of certain …

Deep Learning based discovery of Integrable Systems

2025-03-13 · Shailesh Lal, Suvajit Majumder, Evgeny Sobko

We introduce a novel machine learning based framework for discovering integrable models. Our approach first employs a synchronized ensemble of neural networks to find high-precision numerical solution to the Yang-Baxter …

Deep LearningForm

Avoiding Non-Integrable Beliefs in Expectation Propagation

2026-04-05 · Zilu Zhao, Jichao Chen, Dirk Slock arxiv

Expectation Propagation (EP) is a widely used iterative message-passing algorithm that decomposes a global inference problem into multiple local ones. It approximates marginal distributions as ``beliefs'' using intermedi…

UniMixer: A Unified Architecture for Scaling Laws in Recommendation Systems

2026-04-01 · Mingming Ha, Guanchen Wang, Linxun Chen, Xuan Rao 외 arxiv

In recent years, the scaling laws of recommendation models have attracted increasing attention, which govern the relationship between performance and parameters/FLOPs of recommenders. Currently, there are three mainstrea…

Recommendation Systems

Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration

2026-08-28 · Simeng Sun, Roger Waleffe arxiv

When training Mixture-of-Experts (MoE) language models with expert parallelism, all-to-all token dispatch and combine collectives can consume a substantial fraction of end-to-end training time. In this work, we study com…