paper-with-me

홈 › Papers

FairyFuse: Multiplication-Free LLM Inference on CPUs via Fused Ternary Kernels

2026-04-22 · Fei Zuo, Xiaoyan Xi, Quanyi Zeng, Feiyu Wang, Ho Fai Leung arxiv

Large language models are increasingly deployed on CPU-only platforms where memory bandwidth is the primary bottleneck for autoregressive generation. Weight quantization to four bits or below reduces memory pressure, yet existing systems still dequantize weights and perform floating-point multiplications, limiting the achievable gains. Ternary weights in {-1, 0, +1} provide a more efficient alternative, replacing multiplications with conditional additions, subtractions, or no-ops. While Fairy2i shows that ternary LLMs can match FP16 quality, its runtime does not exploit this structure. We present FairyFuse, an inference system that enables multiplication-free execution on commodity CPUs by fusing the eight real-valued sub-GEMVs of each widely-linear layer into a single AVX-512 loop using masked additions and subtractions, with zero floating-point multiplications. Roofline analysis shows that 16x weight compression shifts memory-bound GEMV toward the compute regime on bandwidth-limited CPUs, yielding a 29.6x kernel speedup while offering little benefit on GPUs. End-to-end, FairyFuse achieves 32.4 tokens per second on a single Intel Xeon 8558P, outperforming llama.cpp Q4_K_M by 1.24x with near-lossless quality (WikiText-2 perplexity 5.52 vs. 5.47 FP16; downstream accuracy 66.0%).

📄 PDF Abstract BibTeX arXiv:2604.20913

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition

2024-01-05 · Adnan Hoque, Less Wright, Chih-Chieh Yang, Mudhakar Srivatsa 외

We propose an implementation of an efficient fused matrix multiplication kernel for W4A16 quantized inference, where we perform dequantization and GEMM in a fused kernel using a SplitK work decomposition. Our implementat…

Litespark Inference For CPUs: Ultra-Fast SIMD Framework for Ternary (1.58-bit) Language Models

2026-05-07 · Nii Osae Osae Dade, Tony Morri, Moinul Hossain Rahat, Sayandip Pal 외 arxiv

Large language models (LLMs) have transformed artificial intelligence, but their computational requirements remain prohibitive for most users. Standard inference demands expensive datacenter GPUs or cloud API access, lea…

FusedMM: A Unified SDDMM-SpMM Kernel for Graph Embedding and Graph Neural Networks

2020-11-07 · Md. Khaledur Rahman, Majedul Haque Sujon, Ariful Azad

We develop a fused matrix multiplication kernel that unifies sampled dense-dense matrix multiplication and sparse-dense matrix multiplication under a single operation called FusedMM. By using user-defined functions, Fuse…

Graph Embedding

Auto-tuning Matrix Multiplication and Convolution for Deep Learning on CPUs

2021-05-21 · NeurIPS 2021 12 · Changbo Chen, Haoyu Chi

Deep learning (DL) compilers have emerged aiming to reduce the gap between abundant, fast-growing DL models and the lag of high performance implementations of these models on diverse hardware devices. In this work, we in…

VibeVoice-ASR-BitNet Technical Report

2026-07-23 · Songchen Xu, Ting Song, Shaohan Huang, Zhiliang Peng 외 arxiv

We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the …