paper-with-me

홈 › Papers

BitLM: Unlocking Multi-Token Language Generation with Bitwise Continuous Diffusion

2026-05-12 · Shaobin Zhuang, Yuang Ai, Jiaming Han, Xiaohui Li, Huaibo Huang, Xiangyu Yue, Xuefeng Hu, Kun Xu, Yali Wang, Hao Chen arxiv

Autoregressive language models generate text one token at a time, yet natural language is inherently structured in multi-token units, including phrases, n-grams, and collocations that carry meaning jointly. This one-token bottleneck limits both the expressiveness of the model during pre-training and its throughput at inference time. Existing remedies such as speculative decoding or diffusion-based language models either leave the underlying bottleneck intact or sacrifice the causal structure essential to language modeling. We propose BitLM, a language model that represents each token as a fixed-length binary code and employs a lightweight diffusion head to denoise multiple tokens in parallel within each block. Crucially, BitLM preserves left-to-right causal attention across blocks while making joint lexical decisions within each block, combining the reliability of autoregressive modeling with the parallelism of iterative refinement. By replacing the large-vocabulary softmax with bitwise denoising, BitLM reframes token generation as iterative commitment in a compact binary space, enabling more efficient pre-training and substantially faster inference without altering the causal foundation that makes language models effective. Our results demonstrate that the one-token-at-a-time paradigm is not a fundamental requirement but an interface choice, and that changing it can yield a stronger and faster language model. We hope BitLM points toward a promising direction for next-generation language model architectures.

📄 PDF Abstract BibTeX arXiv:2605.11577

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Emu3: Next-Token Prediction is All You Need

2024-09-27 · Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun 외

While next-token prediction is considered a promising path towards artificial general intelligence, it has struggled to excel in multimodal tasks, which are still dominated by diffusion models (e.g., Stable Diffusion) an…

AllImage GenerationPrediction+2

Unlocking the Potential of Diffusion Language Models through Template Infilling

2025-10-13 · Junhoo Lee, Seungyeon Kim, Nojun Kwak arxiv

Diffusion Language Models (DLMs) have emerged as a promising alternative to Autoregressive Language Models, yet their inference strategies remain limited to prefix-based prompting inherited from the autoregressive paradi…

Mathematical ReasoningCode Generation

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

2024-03-08 · Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell 외

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions …

1 Image, 2*2 StitchingCode GenerationFS-MEVQAImage Retrieval+8

NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation

2026-01-05 · Huichao Zhang, Liao Qu, Yiheng Liu, Hang Chen 외 arxiv

We present NextFlow, a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image discrete tokens. By leveraging a unified vision representation within a unified autoregressive architect…

Reinforcement LearningVideo GenerationImage Editing

RandAR: Decoder-only Autoregressive Visual Generation in Random Orders

2024-12-02 · CVPR 2025 1 · Ziqi Pang, Tianyuan Zhang, Fujun Luan, Yunze Man 외

We introduce RandAR, a decoder-only visual autoregressive (AR) model capable of generating images in arbitrary token orders. Unlike previous decoder-only AR models that rely on a predefined generation order, RandAR remov…

DecoderInductive Bias