paper-with-me

홈 › Papers

SparseDNN: Fast Sparse Deep Learning Inference on CPUs

2021-01-20 · Ziheng Wang

The last few years have seen gigantic leaps in algorithms and systems to support efficient deep learning inference. Pruning and quantization algorithms can now consistently compress neural networks by an order of magnitude. For a compressed neural network, a multitude of inference frameworks have been designed to maximize the performance of the target hardware. While we find mature support for quantized neural networks in production frameworks such as OpenVINO and MNN, support for pruned sparse neural networks is still lacking. To tackle this challenge, we present SparseDNN, a sparse deep learning inference engine targeting CPUs. We present both kernel-level optimizations with a sparse code generator to accelerate sparse operators and novel network-level optimizations catering to sparse networks. We show that our sparse code generator can achieve significant speedups over state-of-the-art sparse and dense libraries. On end-to-end benchmarks such as Huggingface pruneBERT, SparseDNN achieves up to 5x throughput improvement over dense inference with state-of-the-art OpenVINO. Open source library at: https://github.com/marsupialtail/sparsednn.

📄 PDF Abstract BibTeX arXiv:2101.07948

Code (1)

marsupialtail/sparsednn 공식 구현 pytorch

Tasks

Deep LearningQuantization

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Fast DistilBERT on CPUs

2022-10-27 · Haihao Shen, Ofir Zafrir, Bo Dong, Hengyu Meng 외

Transformer-based language models have become the standard approach to solving natural language processing tasks. However, industry adoption usually requires the maximum throughput to comply with certain latency constrai…

Knowledge DistillationModel CompressionQuantizationQuestion Answering

Real-Time Open-Domain Question Answering with Dense-Sparse Phrase Index

2019-06-13 · ACL 2019 7 · Minjoon Seo, Jinhyuk Lee, Tom Kwiatkowski, Ankur P. Parikh 외

Existing open-domain question answering (QA) models are not suitable for real-time usage because they need to process several long documents on-demand for every input query. In this paper, we introduce the query-agnostic…

GPUOpen-Domain Question AnsweringQuestion Answering

1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs

2024-10-21 · Jinheng Wang, Hansong Zhou, Ting Song, Shaoguang Mao 외

Recent advances in 1-bit Large Language Models (LLMs), such as BitNet and BitNet b1.58, present a promising approach to enhancing the efficiency of LLMs in terms of speed and energy consumption. These developments also e…

Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment

2024-05-06 · Abhinav Agarwalla, Abhay Gupta, Alexandre Marques, Shubhra Pandit 외

Large language models (LLMs) have revolutionized Natural Language Processing (NLP), but their size creates computational bottlenecks. We introduce a novel approach to create accurate, sparse foundational versions of perf…

Arithmetic ReasoningCode GenerationInstruction FollowingQuantization

Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks

2020-01-01 · ICML 2020 1 · Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev 외

Optimizing convolutional neural networks for fast inference has recently become an extremely active area of research. One of the go-to solutions in this context is weight pruning, which aims to reduce computational and m…

image-classificationImage Classification