paper-with-me

홈 › Papers

Multi-DNN Inference of Sparse Models on Edge SoCs

2026-03-10 · Jiawei Luo, Di Wu, Simon Dobson, Blesson Varghese arxiv

Modern edge applications increasingly require multi-DNN inference systems to execute tasks on heterogeneous processors, gaining performance from both concurrent execution and from matching each model to the most suited accelerator. However, existing systems support only a single model (or a few sparse variants) per task, which impedes the efficiency of this matching and results in high Service Level Objective violation rates. We introduce model stitching for multi-DNN inference systems, which creates model variants by recombining subgraphs from sparse models without re-training. We present a demonstrator system, SparseLoom, that shows model stitching can be deployed to SoCs. We show experimentally that SparseLoom reduces SLO violation rates by up to 74%, improves throughput by up to 2.31x, and lowers memory overhead by an average of 28% compared to state-of-the-art multi-DNN inference systems.

📄 PDF Abstract BibTeX arXiv:2603.09642

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Neural Network Inference on Mobile SoCs

2019-08-24 · Siqi Wang, Anuj Pathania, Tulika Mitra

The ever-increasing demand from mobile Machine Learning (ML) applications calls for evermore powerful on-chip computing resources. Mobile devices are empowered with heterogeneous multi-processor Systems-on-Chips (SoCs) t…

CPUGPU

SOCS: Semantically-aware Object Coordinate Space for Category-Level 6D Object Pose Estimation under Large Shape Variations

2023-03-18 · ICCV 2023 1 · Boyan Wan, Yifei Shi, Kai Xu

Most learning-based approaches to category-level 6D pose estimation are design around normalized object coordinate space (NOCS). While being successful, NOCS-based methods become inaccurate and less robust when handling …

6D Pose Estimation6D Pose Estimation using RGBObjectPose Estimation+1

Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices

2025-09-25 · Yilong Li, Shuai Zhang, Yijing Zeng, Hao Zhang 외 arxiv

Large Multimodal Models (LMMs) are inherently modular, comprising vision and audio encoders, a projector, and a language backbone. Yet existing systems execute them monolithically, underutilizing the heterogeneous accele…

Optimizing DNN Inference on Multi-Accelerator SoCs at Training-time

2024-09-27 · Matteo Risso, Alessio Burrello, Daniele Jahier Pagliari

The demand for executing Deep Neural Networks (DNNs) with low latency and minimal power consumption at the edge has led to the development of advanced heterogeneous Systems-on-Chips (SoCs) that incorporate multiple speci…

Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge

2026-05-01 · M. Grailoo, J. Núñez-Yáñez arxiv

Scaling laws for Large Language Models (LLMs) establish that model quality improves with computational scale, yet edge deployment imposes strict constraints on compute, memory, and power. Since General Matrix Multiplicat…