paper-with-me

홈 › Papers

AIPC: Agent-Based Automation for AI Model Deployment with Qualcomm AI Runtime

2026-04-16 · Jianhao Su, Zhanwei Wu, ShengTing Huang, Weidong Feng arxiv

Edge AI model deployment is a multi-stage engineering process involving model conversion, operator compatibility handling, quantization calibration, runtime integration, and accuracy validation. In practice, this workflow is long, failure-prone, and heavily dependent on deployment expertise, particularly when targeting hardware-specific inference runtimes. This technical report presents AIPC (AI Porting Conversion), an AI agent-driven approach for constrained automation of AI model deployment. AIPC decomposes deployment into standardized, verifiable stages and injects deployment-domain knowledge into agent execution through Agent Skills, helper scripts, and a stage-wise validation loop. This design reduces both the expertise barrier and the engineering time required for hardware deployment. Using Qualcomm AI Runtime (QAIRT) as the primary scenario, this report examines automated deployment across representative vision, multimodal, and speech models. In the cases covered here, AIPC can complete deployment from PyTorch to runnable QNN/SNPE inference within 7-20 minutes for structurally regular vision models, with indicative API costs roughly in the range of USD 0.7-10. For more complex models involving less-supported operators, dynamic shapes, or autoregressive decoding structures, fully automated deployment may still require further advances, but AIPC already provides practical support for execution, failure localization, and bounded repair.

📄 PDF Abstract BibTeX arXiv:2604.14661

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Enforcing Company Policy Adherence in Agentic Workflows

2025-07-22 · Naama Zwerdling, David Boaz, Ella Rabinovich, Guy Uziel 외 arxiv

Large Language Model (LLM) agents hold promise for a flexible and scalable alternative to traditional business process automation, but struggle to reliably follow complex company policies. In this study we introduce a de…

A Deployment Case Study in Robotic Apparel Automation: Digital Twin Integration, Interoperability, and Workforce Enablement

2026-06-15 · Gokul Narayanan, Abhiroop Ajith, Jonathan Zornow, Carlos Calle 외 arxiv

Despite steady advances in flexible automation in sectors such as electronics and automotive manufacturing, apparel automation remains challenging because fabrics are deformable and difficult to manipulate with robots. T…

When Bots Take the Bait: Exposing and Mitigating the Emerging Social Engineering Attack in Web Automation Agent

2026-01-12 · Xinyi Wu, Geng Hong, Yueyue Chen, MingXuan Liu 외 arxiv

Web agents, powered by large language models (LLMs), are increasingly deployed to automate complex web interactions. The rise of open-source frameworks (e.g., Browser Use, Skyvern-AI) has accelerated adoption, but also b…

Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs

2025-07-01 · Mohammad Firas Sada, John J. Graham, Elham E Khoda, Mahidhar Tatineni 외 arxiv

This study presents a benchmarking analysis of the Qualcomm Cloud AI 100 Ultra (QAic) accelerator for large language model (LLM) inference, evaluating its energy efficiency (throughput per watt), performance, and hardwar…

Unlocking the Edge deployment and ondevice acceleration of multi-LoRA enabled one-for-all foundational LLM

2026-04-20 · Sravanth Kodavanti, Sowmya Vajrala, Srinivas Miriyala, Utsav Tiwari 외 arxiv

Deploying large language models (LLMs) on smartphones poses significant engineering challenges due to stringent constraints on memory, latency, and runtime flexibility. In this work, we present a hardware-aware framework…