paper-with-me

홈 › Papers

Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs

2025-09-23 · Marcin Chrapek, Marcin Copik, Etienne Mettaz, Torsten Hoefler arxiv

Large Language Models (LLMs) are increasingly deployed on converged Cloud and High-Performance Computing (HPC) infrastructure. However, as LLMs handle confidential inputs and are fine-tuned on costly, proprietary datasets, their heightened security requirements slow adoption in privacy-sensitive sectors such as healthcare and finance. We investigate methods to address this gap and propose Trusted Execution Environments (TEEs) as a solution for securing end-to-end LLM inference. We validate their practicality by evaluating these compute-intensive workloads entirely within CPU and GPU TEEs. On the CPU side, we conduct an in-depth study running full Llama2 inference pipelines (7B, 13B, 70B) inside Intel's TDX and SGX, accelerated by Advanced Matrix Extensions (AMX). We derive 12 insights, including that across various data types, batch sizes, and input lengths, CPU TEEs impose under 10% throughput and 20% latency overheads, further reduced by AMX. We run LLM inference on NVIDIA H100 Confidential Compute GPUs, contextualizing our CPU findings and observing throughput penalties of 4-8% that diminish as batch and input sizes grow. By comparing performance, cost, and security trade-offs, we show how CPU TEEs can be more cost-effective or secure than their GPU counterparts. To our knowledge, our work is the first to comprehensively demonstrate the performance and practicality of modern TEEs across both CPUs and GPUs for enabling confidential LLMs (cLLMs).

📄 PDF Abstract BibTeX arXiv:2509.18886

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Confidential and Efficient LLM Inference with Dual Privacy Protection

2025-09-11 · Honglan Yu, Yibin Wang, Feifei Dai, Dong Liu 외 arxiv

CPU-based trusted execution environments (TEEs) and differential privacy (DP) have gained wide applications for private inference. Due to high inference latency in TEEs, researchers use partition-based approaches that of…

Machine Learning with Confidential Computing: A Systematization of Knowledge

2022-08-22 · Fan Mo, Zahra Tarkhani, Hamed Haddadi

Privacy and security challenges in Machine Learning (ML) have become increasingly severe, along with ML's pervasive development and the recent demonstration of large attack surfaces. As a mature system-oriented approach,…

Dstack: A Zero Trust Framework for Confidential Containers

2025-09-15 · Shunfan Zhou, Kevin Wang, Hang Yin arxiv

Web3 applications require execution platforms that maintain confidentiality and integrity without relying on centralized trust authorities. While Trusted Execution Environments (TEEs) offer promising capabilities for con…

GuardNN: Secure Accelerator Architecture for Privacy-Preserving Deep Learning

2020-08-26 · Weizhe Hua, Muhammad Umar, Zhiru Zhang, G. Edward Suh

This paper proposes GuardNN, a secure DNN accelerator that provides hardware-based protection for user data and model parameters even in an untrusted environment. GuardNN shows that the architecture and protection can be…

Deep LearningPrivacy PreservingPrivacy Preserving Deep Learning

Securing Transformer-based AI Execution via Unified TEEs and Crypto-protected Accelerators

2025-07-04 · Jiaqi Xue, Yifei Zhao, Mengxin Zheng, Fan Yao 외 arxiv

Recent advances in Transformer models, e.g., large language models (LLMs), have brought tremendous breakthroughs in various artificial intelligence (AI) tasks, leading to their wide applications in many security-critical…