paper-with-me

홈 › Papers

Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge

2026-05-01 · M. Grailoo, J. Núñez-Yáñez arxiv

Scaling laws for Large Language Models (LLMs) establish that model quality improves with computational scale, yet edge deployment imposes strict constraints on compute, memory, and power. Since General Matrix Multiplication (GEMM) accounts for up to 90% of inference time, efficient GEMM acceleration is critical for edge AI. The Adaptive Intelligent Engines available in the AMD Versal adaptive SoCs are well suited for this task, but existing state-of-the-art (SOTA) frameworks maximize performance through spatial scaling, distributing workloads across hundreds of cores -- an approach that fails on resource-limited edge SoCs due to physical implementation failures, bandwidth saturation, and excessive resource consumption. We propose Tempus, a Resource-Invariant Temporal GEMM framework for the AMD Versal AI Edge SoC. Rather than expanding hardware resources with matrix size, Tempus employs a fixed compute block of 16 AIE-ML cores, achieving scalability through iterative graph execution and algorithmic data tiling and replication in the Programmable Logic. High-speed cascade streaming ensures low-latency partial sum reduction at Initiation Interval (II) of 1, while a deadlock-free DATAFLOW protocol maximizes transfer-compute overlap and PLIO reuse. Evaluated on GEMM workloads, Tempus achieves 607 GOPS at 10.677 W total on-chip power. By characterizing system-level efficiency through the Platform-Aware Utility (PAU) metric, we prove that Tempus achieves a 211.2x higher prominence factor than the leading spatial SOTA (ARIES). Furthermore, the framework maintains a 0.00% utilization of URAM/DSP, yielding 22.0x core frugality, 7.1x power frugality, and a 6.3x reduction in I/O demand, establishing a sustainable, scalable foundation for edge LLM inference.

📄 PDF Abstract BibTeX arXiv:2605.00536

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs

2024-12-25 · Prabhu Vellaisamy, Harideep Nair, Thomas Kang, Yichen Ni 외

The increasing complexity of deep neural networks (DNNs) poses significant challenges for edge inference deployment due to resource and power constraints of edge devices. Recent works on unary-based matrix multiplication…

TempusBench: An Evaluation Framework for Time-Series Forecasting

2026-04-13 · Denizalp Goktas, Gerardo Riaño-Briceño, Alif Abdullah, Aryan Nair 외 arxiv

Foundation models have transformed natural language processing and computer vision, and a rapidly growing literature on time-series foundation models (TSFMs) seeks to replicate this success in forecasting. While recent o…

Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy

2025-09-16 · Nadim Barakat, William Lotter arxiv

Diabetic retinopathy (DR) is a leading cause of blindness worldwide, and AI systems can expand access to fundus photography screening. Current FDA-cleared systems primarily provide binary referral outputs, where this min…

Invariant-based Robust Weights Watermark for Large Language Models

2025-07-11 · Qingxiao Guo, Xinjie Zhu, Yilong Ma, Hui Jin 외 arxiv

Watermarking technology has gained significant attention due to the increasing importance of intellectual property (IP) rights, particularly with the growing deployment of large language models (LLMs) on billions resourc…

MedGemma 1.5 Technical Report

2026-04-06 · Andrew Sellergren, Chufan Gao, Fereshteh Mahvar, Timo Kohlberger 외 arxiv

We introduce MedGemma 1.5 4B, the latest model in the MedGemma collection. MedGemma 1.5 expands on MedGemma 1 by integrating additional capabilities: high-dimensional medical imaging (CT/MRI volumes and histopathology wh…

Information ExtractionClinical Knowledge