paper-with-me

홈 › Papers

ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM

2026-07-16 · Hyunwoo Oh, Suyeon Jang, Hanning Chen, Sanggeon Yun, Ryozo Masukawa, Mohsen Imani arxiv

Low-bit GEMM is increasingly central to efficient ML inference, yet very-low-bit execution remains a poor fit for conventional CPUs. Practical deployment spans fragmented regimes-from 1/2/4-bit weights to varying activation precision-whose feasibility, reuse opportunity, and support cost differ under fixed SIMD and register-file budgets, making lightweight CPU support selection a first-class design problem. We present ExaGEMM, a workload-aware codesign and exploration framework for CPU-native low-bit GEMM via register-resident LUT execution. The key insight is that existing SIMD datapaths already cover table generation and accumulation; the only new hardware is an in-register select/feed mechanism with explicitly modeled cost. ExaGEMM co-explores parameterized kernels and lightweight SIMD ISA support using analytical models of register feasibility, compute cost, memory traffic, and hardware overhead, pruning the candidate space by 99.2% before simulation. It then identifies non-dominated support points and generates ISA specs, gem5 patches, and GEMM kernels for validation. Across representative ML models and CPU targets, ExaGEMM improves latency by 13.29x over software-only baselines, while showing that workload-aware frontier selection is especially important for mixed-precision LLM workloads.

📄 PDF Abstract BibTeX arXiv:2607.14622

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mimicking associative learning of rats via a neuromorphic robot in open field maze using spatial cell models

2025-08-25 · Tianze Liu, Md Abu Bakr Siddique, Hongyu An arxiv

Data-driven Artificial Intelligence (AI) approaches have exhibited remarkable prowess across various cognitive tasks using extensive training data. However, the reliance on large datasets and neural networks presents cha…

On the Relationship Between Variational Inference and Auto-Associative Memory

2022-10-14 · Louis Annabi, Alexandre Pitti, Mathias Quoy

In this article, we propose a variational inference formulation of auto-associative memories, allowing us to combine perceptual inference and memory retrieval into the same mathematical framework. In this formulation, th…

RetrievalVariational Inference

CoAT: Chain-of-Associated-Thoughts Framework for Enhancing Large Language Models Reasoning

2025-02-04 · Jianfeng Pan, Senyou Deng, Shaomang Huang

Research on LLM technologies is rapidly emerging, with most of them employing a 'fast thinking' approach to inference. Most LLMs generate the final result based solely on a single query and LLM's reasoning capabilities. …

Towards Modeling the Interaction of Spatial-Associative Neural Network Representations for Multisensory Perception

2018-07-13 · German I. Parisi, Jonathan Tong, Pablo Barros, Brigitte Röder 외

Our daily perceptual experience is driven by different neural mechanisms that yield multisensory interaction as the interplay between exogenous stimuli and endogenous expectations. While the interaction of multisensory c…

Causal Inference

Firing Rate Models as Associative Memory: Excitatory-Inhibitory Balance for Robust Retrieval

2024-11-11 · Simone Betteti, Giacomo Baggio, Francesco Bullo, Sandro Zampieri

Firing rate models are dynamical systems widely used in applied and theoretical neuroscience to describe local cortical dynamics in neuronal populations. By providing a macroscopic perspective of neuronal activity, these…

Retrieval