paper-with-me

홈 › Papers

How Much Progress Has There Been in NVIDIA Datacenter GPUs?

2026-01-27 · Emanuele Del Sozzo, Martin Fleming, Kenneth Flamm, Neil Thompson arxiv

As the role of modern Graphics Processing Units (GPUs) becomes increasingly essential for several computing tasks, analyzing their past and current progress is paramount for determining future constraints on scientific research. This is particularly compelling in the Artificial Intelligence (AI) domain, where rapid technological advancements and fierce global competition have led the United States to recently implement export control regulations limiting international access to advanced AI chips. Consequently, this paper examines technical progress in NVIDIA datacenter GPUs from the mid-2000s through 2025. Our main results identify doubling times of 1.43 and 1.67 years for FP16 and FP32 dense operations, while FP64 doubling times range from 2.05 to 3.79 years. Off-chip memory size and bandwidth have grown at slower rates than computing performance, doubling every 3.29 to 3.41 years, whereas the release prices and power consumption roughly doubled every 5.03 and 15 years, respectively. Moreover, our cross-vendor comparison of the top-performing GPUs per year shows that NVIDIA's performance advantage is narrowing, but not enough to compel a major market shift. Finally, we quantify the potential implications of current U.S. export control regulations and the consequent performance gaps, which the recently proposed policy changes could shrink from 23.6X to 3.54X.

📄 PDF Abstract BibTeX arXiv:2601.20115

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Implementing Reinforcement Learning Datacenter Congestion Control in NVIDIA NICs

2022-07-05 · Benjamin Fuhrer, Yuval Shpigelman, Chen Tessler, Shie Mannor 외

As communication protocols evolve, datacenter network utilization increases. As a result, congestion is more frequent, causing higher latency and packet loss. Combined with the increasing complexity of workloads, manual …

Fairnessreinforcement-learningReinforcement LearningReinforcement Learning (RL)

In-Datacenter Performance Analysis of a Tensor Processing Unit

2017-04-16 · Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson 외

Many architects believe that major improvements in cost-energy-performance must now come from domain-specific hardware. This paper evaluates a custom ASIC---called a Tensor Processing Unit (TPU)---deployed in datacenters…

CPUGPU

Searching for Fast Model Families on Datacenter Accelerators

2021-02-10 · CVPR 2021 1 · Sheng Li, Mingxing Tan, Ruoming Pang, Andrew Li 외

Neural Architecture Search (NAS), together with model scaling, has shown remarkable progress in designing high accuracy and fast convolutional architecture families. However, as neither NAS nor model scaling considers su…

modelNeural Architecture Search

MOSAIC: A Multi-Objective Optimization Framework for Sustainable Datacenter Management

2023-11-14 · Sirui Qi, Dejan Milojicic, Cullen Bash, Sudeep Pasricha

In recent years, cloud service providers have been building and hosting datacenters across multiple geographical locations to provide robust services. However, the geographical distribution of datacenters introduces grow…

Management

DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation

2022-09-22 · Seongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee 외

Transformer is a deep learning language model widely used for natural language processing (NLP) services in datacenters. Among transformer models, Generative Pre-trained Transformer (GPT) has achieved remarkable performa…

GPULanguage ModellingText Generation