paper-with-me

Papers

Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception

2024-11-24 · Mohanad Odema, Luke Chen, Hyoukjun Kwon, Mohammad Abdullah Al Faruque

We study the application of emerging chiplet-based Neural Processing Units to accelerate vehicular AI perception workloads in constrained automotive settings. The motivation stems from how chiplets technology is becoming integral to emerging vehicular architectures, providing a cost-effective trade-off between performance, modularity, and customization; and from perception models being the most computationally demanding workloads in a autonomous driving system. Using the Tesla Autopilot perception pipeline as a case study, we first breakdown its constituent models and profile their performance on different chiplet accelerators. From the insights, we propose a novel scheduling strategy to efficiently deploy perception workloads on multi-chip AI accelerators. Our experiments using a standard DNN performance simulator, MAESTRO, show our approach realizes 82% and 2.8x increase in throughput and processing engines utilization compared to monolithic accelerator designs.

📄 PDF Abstract BibTeX arXiv:2411.16007

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingScheduling

Similar Papers 제목 키워드 기반

Chiplet-Based RISC-V SoC with Modular AI Acceleration

2025-09-22 · Suhas Suresh Bharadwaj, Prerana Ramkumar arxiv

Achieving high performance, energy efficiency, and cost-effectiveness while maintaining architectural flexibility is a critical challenge in the development and deployment of edge AI devices. Monolithic SoC designs strug…

HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads

2026-08-19 · Jiahao Lin, Alish Kanani, Sangwan Lee, Jaehyun Park 외 arxiv

Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communication patterns complicate efficient hardware acceleration. Chiplet-based architecture…

Sangam: Chiplet-Based DRAM-PIM Accelerator with CXL Integration for LLM Inferencing

2025-11-15 · Khyati Kiyawat, Zhenxing Fan, Yasas Seneviratne, Morteza Baradaran 외 arxiv

Large Language Models (LLMs) are becoming increasingly data-intensive due to growing model sizes, and they are becoming memory-bound as the context length and, consequently, the key-value (KV) cache size increase. Infere…

Machine Learning Accelerators in 2.5D Chiplet Platforms with Silicon Photonics

2023-01-28 · Febin Sunny, Ebadollah Taheri, Mahdi Nikdast, Sudeep Pasricha

Domain-specific machine learning (ML) accelerators such as Google's TPU and Apple's Neural Engine now dominate CPUs and GPUs for energy-efficient ML processing. However, the evolution of electronic accelerators is facing…

Chiplet Placement Order Exploration Based on Learning to Rank with Graph Representation

2024-04-07 · Zhihui Deng, Yuanyuan Duan, Leilai Shao, Xiaolei Zhu

Chiplet-based systems, integrating various silicon dies manufactured at different integrated circuit technology nodes on a carrier interposer, have garnered significant attention in recent years due to their cost-effecti…

Learning-To-Rankreinforcement-learningReinforcement Learning