paper-with-me

홈 › Papers

TriGen: NPU Architecture for End-to-End Acceleration of Large Language Models based on SW-HW Co-Design

2026-02-13 · Jonghun Lee, Junghoon Lee, Hyeonjin Kim, Seoho Jeon, Jisup Yoon, Hyunbin Park, Meejeong Park, Heonjae Ha arxiv

Recent studies have extensively explored NPU architectures for accelerating AI inference in on-device environments, which are inherently resource-constrained. Meanwhile, transformer-based large language models (LLMs) have become dominant, with rapidly increasing model sizes but low degree of parameter reuse compared to conventional CNNs, making end-to-end execution on resource-limited devices extremely challenging. To address these challenges, we propose TriGen, a novel NPU architecture tailored for resource-constrained environments through software-hardware co-design. Firstly, TriGen adopts low-precision computation using microscaling (MX) to enable additional optimization opportunities while preserving accuracy, and resolves the issues that arise by employing such precision. Secondly, to jointly optimize both nonlinear and linear operations, TriGen eliminates the need for specialized hardware for essential nonlinear operations by using fast and accurate LUT, thereby maximizing performance gains and reducing hardware-cost in on-device environments, and finally, by taking practical hardware constraints into account, further employs scheduling techniques to maximize computational utilization even under limited on-chip memory capacity. We evaluate the performance of TriGen on various LLMs and show that TriGen achieves an average 2.73x performance speedup and 52% less memory transfer over the baseline NPU design with negligible accuracy loss.

📄 PDF Abstract BibTeX arXiv:2602.12962

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pruning Algorithms for Low-Dimensional Non-metric k-NN Search: A Case Study

2019-10-08 · Leonid Boytsov, Eric Nyberg

We focus on low-dimensional non-metric search, where tree-based approaches permit efficient and accurate retrieval while having short indexing time. These methods rely on space partitioning and require a pruning rule to …

Retrieval

AttriGen: Automated Multi-Attribute Annotation for Blood Cell Datasets

2025-09-30 · Walid Houmaidi, Youssef Sabiri, Fatima Zahra Iguenfer, Amine Abouaomar arxiv

We introduce AttriGen, a novel framework for automated, fine-grained multi-attribute annotation in computer vision, with a particular focus on cell microscopy where multi-attribute classification remains underrepresented…

Harnessing Diet and Gene Expression Insights through a Centralized Nutrigenomics Database to Improve Public Health

2025-06-23 · Fahmida Hai, Shriya Samudrala, Ijeoma Ezengwa, Rubayat Khan 외

Nutrigenomics is an emerging field that explores the intricate interaction between genes and diet. This study aimed to develop a comprehensive database to help clinicians and patients understand the connections between g…

UniTriGen: Unified Triplet Generation of Aligned Visible-Infrared-Label for Few-Shot RGB-T Semantic Segmentation

2026-05-14 · Ping Zhou, Haoyu Wang, Mengmeng Zheng, Lei Zhang 외 arxiv

RGB-T semantic segmentation requires strictly aligned VIS-IR-Label triplets; however, such aligned triplet data are often scarce in real-world scenarios. Existing generative augmentation methods usually adopt cascaded ge…

Semantic Segmentation

NutriGen: Personalized Meal Plan Generator Leveraging Large Language Models to Enhance Dietary and Nutritional Adherence

2025-02-28 · Saman Khamesian, Asiful Arefeen, Stephanie M. Carpenter, Hassan Ghasemzadeh

Maintaining a balanced diet is essential for overall health, yet many individuals struggle with meal planning due to nutritional complexity, time constraints, and lack of dietary knowledge. Personalized food recommendati…

NutritionPrompt EngineeringRecommendation Systems