paper-with-me

홈 › Papers

How to Reach Real-Time AI on Consumer Devices? Solutions for Programmable and Custom Architectures

2021-06-21 · Stylianos I. Venieris, Ioannis Panopoulos, Ilias Leontiadis, Iakovos S. Venieris

The unprecedented performance of deep neural networks (DNNs) has led to large strides in various Artificial Intelligence (AI) inference tasks, such as object and speech recognition. Nevertheless, deploying such AI models across commodity devices faces significant challenges: large computational cost, multiple performance objectives, hardware heterogeneity and a common need for high accuracy, together pose critical problems to the deployment of DNNs across the various embedded and mobile devices in the wild. As such, we have yet to witness the mainstream usage of state-of-the-art deep learning algorithms across consumer devices. In this paper, we provide preliminary answers to this potentially game-changing question by presenting an array of design techniques for efficient AI systems. We start by examining the major roadblocks when targeting both programmable processors and custom accelerators. Then, we present diverse methods for achieving real-time performance following a cross-stack approach. These span model-, system- and hardware-level techniques, and their combination. Our findings provide illustrative examples of AI systems that do not overburden mobile hardware, while also indicating how they can improve inference accuracy. Moreover, we showcase how custom ASIC- and FPGA-based accelerators can be an enabling factor for next-generation AI applications, such as multi-DNN systems. Collectively, these results highlight the critical need for further exploration as to how the various cross-stack solutions can be best combined in order to bring the latest advances in deep learning close to users, in a robust and efficient manner.

📄 PDF Abstract BibTeX arXiv:2106.15021

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

DENALI: A Dataset Enabling Non-Line-of-Sight Spatial Reasoning with Low-Cost LiDARs

2026-04-17 · Nikhil Behari, Diego Rivero, Luke Apostolides, Suman Ghosh 외 arxiv

Consumer LiDARs in mobile devices and robots typically output a single depth value per pixel. Yet internally, they record full time-resolved histograms containing direct and multi-bounce light returns; these multi-bounce…

Spatial Reasoning

MLVC: Multi-platform Learned Video Codec for Real-World Deployment

2026-06-26 · Tanel Pärnamaa, Martin Lumiste, Ardi Loot, Evgenii Indenbom 외 arxiv

Neural video codecs have surpassed classical codecs in coding efficiency but remain impractical for deployment due to cross-platform incompatibility and high computational cost. Existing quantization-based solutions fail…

A Review of Indoor Millimeter Wave Device-based Localization and Device-free Sensing Technologies and Applications

2021-12-10 · Anish Shastri, Neharika Valecha, Enver Bashirov, Harsh Tataria 외

The commercial availability of low-cost millimeter wave (mmWave) communication and radar devices is starting to improve the penetration of such technologies in consumer markets, paving the way for large-scale and dense d…

SwiftVR: Real-Time One-Step Generative Video Restoration

2026-06-08 · Jiaqi Yan, Xiangyu Chen, Xinlin Zhong, Haibin Huang 외 arxiv

Real-time video restoration (VR) for live streams requires high-resolution outputs under strict per-frame latency constraints. Existing one-step diffusion-based VR models remain difficult to deploy on consumer-grade GPUs…

Video Restoration

ConsumerBench: Benchmarking Generative AI Applications on End-User Devices

2025-06-21 · Yile Gu, Rohan Kadekodi, Hoang Nguyen, Keisuke Kamahori 외

The recent shift in Generative AI (GenAI) applications from cloud-only environments to end-user devices introduces new challenges in resource management, system efficiency, and user experience. This paper presents Consum…

BenchmarkingCPUGPUScheduling