paper-with-me

Papers

Efficient Large Language Models with Zero-Shot Adjustable Acceleration

2025-09-01 · Sajjad Kachuee, Mohammad Sharifkhani arxiv

Using Large Language Models (LLMs) in real-world applications presents significant challenges, particularly in balancing computational efficiency with model performance. Optimizing acceleration after fine-tuning and during inference is critical for building efficient architectures. This paper introduces Zero-Shot Adjustable Acceleration, a novel training and inference method that dynamically adjusts hardware utilization during inference without requiring additional fine-tuning. The proposed approach is applied to recent LLMs and evaluated across multiple classification and text generation tasks. Experimental results demonstrate that the method supports a wide range of zero-shot acceleration and achieves up to 11x speedup compared to the baseline.

📄 PDF Abstract BibTeX arXiv:2509.01190

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyText Generation

Similar Papers 제목 키워드 기반

Retcon -- a Prompt-Based Technique for Precise Control of LLMs in Conversations

2026-02-09 · David Kogan, Sam Nguyen, Masanori Suzuki, Feiyang Chen arxiv

Recent advances in Large Language Models (LLMs) allow agents to execute complex natural language tasks. Many LLM applications, such as support agents, teaching assistants, and interactive bots, involve multi-turn convers…

SubZero: Subspace Zero-Shot MRI Reconstruction

2023-11-28 · Heng Yu, Yamin Arefeen, Berkin Bilgic

Recently introduced zero-shot self-supervised learning (ZS-SSL) has shown potential in accelerated MRI in a scan-specific scenario, which enabled high-quality reconstructions without access to a large training dataset. Z…

MRI ReconstructionSelf-Supervised Learning

Agent Instructs Large Language Models to be General Zero-Shot Reasoners

2023-10-05 · Nicholas Crispino, Kyle Montgomery, Fankun Zeng, Dawn Song 외

We introduce a method to improve the zero-shot reasoning abilities of large language models on general language understanding tasks. Specifically, we build an autonomous agent to instruct the reasoning process of large l…

ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks

2023-12-14 · Xiaoxia Wu, Haojun Xia, Stephen Youn, Zhen Zheng 외

This study examines 4-bit quantization methods like GPTQ in large language models (LLMs), highlighting GPTQ's overfitting and limited enhancement in Zero-Shot tasks. While prior works merely focusing on zero-shot measure…

Abstractive Text SummarizationCode GenerationQuantization

NLCG-Net: A Model-Based Zero-Shot Learning Framework for Undersampled Quantitative MRI Reconstruction

2024-01-22 · Xinrui Jiang, Yohan Jun, Jaejin Cho, Mengze Gao 외

Typical quantitative MRI (qMRI) methods estimate parameter maps after image reconstructing, which is prone to biases and error propagation. We propose a Nonlinear Conjugate Gradient (NLCG) optimizer for model-based T2/T1…

MRI ReconstructionQuantitative MRIZero-Shot Learning