paper-with-me

홈 › Papers

FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference

2021-01-13 · Daya Khudia, Jianyu Huang, Protonu Basu, Summer Deng, Haixin Liu, Jongsoo Park, Mikhail Smelyanskiy

Deep learning models typically use single-precision (FP32) floating point data types for representing activations and weights, but a slew of recent research work has shown that computations with reduced-precision data types (FP16, 16-bit integers, 8-bit integers or even 4- or 2-bit integers) are enough to achieve same accuracy as FP32 and are much more efficient. Therefore, we designed fbgemm, a high-performance kernel library, from ground up to perform high-performance quantized inference on current generation CPUs. fbgemm achieves efficiency by fusing common quantization operations with a high-performance gemm implementation and by shape- and size-specific kernel code generation at runtime. The library has been deployed at Facebook, where it delivers greater than 2x performance gains with respect to our current production baseline.

📄 PDF Abstract BibTeX arXiv:2101.05615

Code (1)

pytorch/fbgemm 공식 구현 pytorch

Tasks

Code GenerationDeep LearningQuantizationVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

Debunking the CUDA Myth Towards GPU-based AI Systems

2024-12-31 · Yunjae Lee, Juntaek Lim, Jehyeon Bang, Eunyeong Cho 외

This paper presents a comprehensive evaluation of Intel Gaudi NPUs as an alternative to NVIDIA GPUs, which is currently the de facto standard in AI system design. First, we create a suite of microbenchmarks to compare In…

GPU

BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling

2026-02-02 · Zisheng Ye, Xiaoyu He, Maoyuan Song, Guoliang Qiu 외 arxiv

As the performance gains from accelerating quantized matrix multiplication plateau, the softmax operation becomes the critical bottleneck in Transformer inference. This bottleneck stems from two hardware limitations: (1)…

KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache

2025-05-18 · Fei Li, Song Liu, Weiguo Wu, Shiqiang Nie 외

The high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quantization can effectively alleviate the me…

Quantization

Inference for High-Dimensional Sparse Spectral Precision Matrices

2026-06-06 · Navonil Deb, Younghoon Kim, Sumanta Basu arxiv

Gaussian graphical models in the spectral domain offer a principled approach for recovering conditional dependence structures in stationary high-dimensional time series. Inference on the spectral precision matrix at a fi…

PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction

2025-03-21 · Ting Sun, Cheng Cui, Yuning Du, Yi Liu

Document layout analysis is a critical preprocessing step in document intelligence, enabling the detection and localization of structural elements such as titles, text blocks, tables, and formulas. Despite its importance…

CPUDocument Layout AnalysisGPU