paper-with-me

홈 › Papers

When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization

2024-11-08 · Jacob Nielsen, Lukas Galke, Peter Schneider-Kamp

Contemporary machine learning models, such as language models, are powerful, but come with immense resource requirements both at training and inference time. It has been shown that decoder-only language models can be trained to a competitive state with ternary weights (1.58 bits per weight), facilitating efficient inference. Here, we start our exploration with non-transformer model architectures, investigating 1.58-bit training for multi-layer perceptrons and graph neural networks. Then, we explore 1.58-bit training in other transformer-based language models, namely encoder-only and encoder-decoder models. Our results show that in all of these settings, 1.58-bit training is on par with or sometimes even better than the standard 32/16-bit models.

📄 PDF Abstract BibTeX arXiv:2411.05882

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderQuantization

Similar Papers 제목 키워드 기반

Bitnet.cpp: Efficient Edge Inference for Ternary LLMs

2025-02-17 · Jinheng Wang, Hansong Zhou, Ting Song, Shijie Cao 외

The advent of 1-bit large language models (LLMs), led by BitNet b1.58, has spurred interest in ternary LLMs. Despite this, research and practical applications focusing on efficient edge inference for ternary LLMs remain …

Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity

2026-03-05 · Di Zhang, Xun Wu, Shaohan Huang, Yudong Wang 외 arxiv

Semi-structured N:M sparsity and low-bit quantization (e.g., 1.58-bit BitNet) are two promising approaches for improving the efficiency of large language models (LLMs), yet they have largely been studied in isolation. In…

BitNet: Scaling 1-bit Transformers for Large Language Models

2023-10-17 · Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang 외

The increasing size of large language models has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. In this work, we introduce BitNet, a scalable and stable 1-b…

Language ModelingLanguage ModellingQuantization

MAGNET: Autonomous Expert Model Generation via Decentralized Autoresearch and BitNet Training

2026-03-26 · Yongwan Kim, Sungchul Park arxiv

We present MAGNET (Model Autonomously Growing Network), a decentralized system for autonomous generation, training, and serving of domain-expert language models across commodity hardware. MAGNET integrates four component…

Hyperparameter Optimization

NativeTernary: A Self-Delimiting Binary Encoding with Unary Run-Length Hierarchy Markers for Ternary Neural Network Weights, Structured Data, and General Computing Infrastructure

2026-04-03 · Maharshi Savdhariya arxiv

BitNet b1.58 (Ma et al., 2024) demonstrates that large language models can operate entirely on ternary weights {-1, 0, +1}, yet no native binary wire format exists for such models. NativeTernary closes this gap. Benchmar…