paper-with-me

Papers

NVCiM-PT: An NVCiM-assisted Prompt Tuning Framework for Edge LLMs

2024-11-12 · Ruiyang Qin, Pengyu Ren, Zheyu Yan, Liu Liu, Dancheng Liu, Amir Nassereldine, JinJun Xiong, Kai Ni, Sharon Hu, Yiyu Shi

Large Language Models (LLMs) deployed on edge devices, known as edge LLMs, need to continuously fine-tune their model parameters from user-generated data under limited resource constraints. However, most existing learning methods are not applicable for edge LLMs because of their reliance on high resources and low learning capacity. Prompt tuning (PT) has recently emerged as an effective fine-tuning method for edge LLMs by only modifying a small portion of LLM parameters, but it suffers from user domain shifts, resulting in repetitive training and losing resource efficiency. Conventional techniques to address domain shift issues often involve complex neural networks and sophisticated training, which are incompatible for PT for edge LLMs. Therefore, an open research question is how to address domain shift issues for edge LLMs with limited resources. In this paper, we propose a prompt tuning framework for edge LLMs, exploiting the benefits offered by non-volatile computing-in-memory (NVCiM) architectures. We introduce a novel NVCiM-assisted PT framework, where we narrow down the core operations to matrix-matrix multiplication, which can then be accelerated by performing in-situ computation on NVCiM. To the best of our knowledge, this is the first work employing NVCiM to improve the edge LLM PT performance.

📄 PDF Abstract BibTeX arXiv:2411.08244

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TSB: Tiny Shared Block for Efficient DNN Deployment on NVCIM Accelerators

2024-05-08 · Yifan Qin, Zheyu Yan, Zixuan Pan, Wujie Wen 외

Compute-in-memory (CIM) accelerators using non-volatile memory (NVM) devices offer promising solutions for energy-efficient and low-latency Deep Neural Network (DNN) inference execution. However, practical deployment is …

E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference

2026-05-20 · Ankit Kumar Tenwar, Mukul Lokhande, Santosh Kumar Vishvakarma arxiv

This work presents E-ReCON, a 16 Kb energy and resource-efficient digital compute-in-memory (DCIM) macro based on a compact 3T1R ReRAM bitcell for edge-AI inference. The proposed bitcell occupies only 0.85 um^2 and suppo…

Negative Feedback Training: A Novel Concept to Improve Robustness of NVCIM DNN Accelerators

2023-05-23 · Yifan Qin, Zheyu Yan, Wujie Wen, Xiaobo Sharon Hu 외

Compute-in-memory (CIM) accelerators built upon non-volatile memory (NVM) devices excel in energy efficiency and latency when performing Deep Neural Network (DNN) inference, thanks to their in-situ data processing capabi…

Improving Realistic Worst-Case Performance of NVCiM DNN Accelerators through Training with Right-Censored Gaussian Noise

2023-07-29 · Zheyu Yan, Yifan Qin, Wujie Wen, Xiaobo Sharon Hu 외

Compute-in-Memory (CiM), built upon non-volatile memory (NVM) devices, is promising for accelerating deep neural networks (DNNs) owing to its in-situ data processing capability and superior energy efficiency. Unfortunate…

Self-Driving Cars

Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis

2025-08-22 · Teo Susnjak arxiv

Large language models (LLMs) offer significant potential to accelerate systematic literature reviews (SLRs), yet current approaches often rely on brittle, manually crafted prompts that compromise reliability and reproduc…