paper-with-me

홈 › Papers

Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks

2025-07-10 · Joyeeta Datta, Niclas Doll, Qusai Ramadan, Zeyd Boukhers arxiv

Large Language Models (LLMs) have demonstrated outstanding performance across a range of NLP tasks, however, their computational demands hinder their deployment in real-world, resource-constrained environments. This work investigates the extent to which LLMs can be compressed using Knowledge Distillation (KD) while maintaining strong performance on Question Answering (QA) tasks. We evaluate student models distilled from the Pythia and Qwen2.5 families on two QA benchmarks, SQuAD and MLQA, under zero-shot and one-shot prompting conditions. Results show that student models retain over 90% of their teacher models' performance while reducing parameter counts by up to 57.1%. Furthermore, one-shot prompting yields additional performance gains over zero-shot setups for both model families. These findings underscore the trade-off between model efficiency and task performance, demonstrating that KD, combined with minimal prompting, can yield compact yet capable QA systems suitable for resource-constrained applications.

📄 PDF Abstract BibTeX arXiv:2507.07630

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationQuestion AnsweringModel Compression

Similar Papers 제목 키워드 기반

Exploring the Limits of Simple Learners in Knowledge Distillation for Document Classification with DocBERT

2020-07-01 · WS 2020 7 · Ashutosh Adhikari, Achyudh Ram, Raphael Tang, William L. Hamilton 외

Fine-tuned variants of BERT are able to achieve state-of-the-art accuracy on many natural language processing tasks, although at significant computational costs. In this paper, we verify BERT{'}s effectiveness for docume…

Document ClassificationGeneral ClassificationKnowledge DistillationModel Compression

A Functional Perspective on Knowledge Distillation in Neural Networks

2025-10-14 · Israel Mason-Williams, Gabryel Mason-Williams, Helen Yannakoudakis arxiv

Knowledge distillation is considered a compression mechanism when judged on the resulting student's accuracy and loss, yet its functional impact is poorly understood. We quantify the compression capacity of knowledge dis…

Knowledge Distillation

Minimizing PLM-Based Few-Shot Intent Detectors

2024-07-13 · Haode Zhang, Albert Y. S. Lam, Xiao-Ming Wu

Recent research has demonstrated the feasibility of training efficient intent detectors based on pre-trained language model~(PLM) with limited labeled data. However, deploying these detectors in resource-constrained envi…

Data AugmentationKnowledge DistillationLanguage ModelingLanguage Modelling+1

Parallel Blockwise Knowledge Distillation for Deep Neural Network Compression

2020-12-05 · Cody Blakeney, Xiaomin Li, Yan Yan, Ziliang Zong

Deep neural networks (DNNs) have been extremely successful in solving many challenging AI tasks in natural language processing, speech recognition, and computer vision nowadays. However, DNNs are typically computation in…

Knowledge DistillationNeural Network CompressionQuantizationspeech-recognition+1

Microdosing: Knowledge Distillation for GAN based Compression

2022-01-07 · Leonhard Helminger, Roberto Azevedo, Abdelaziz Djelouah, Markus Gross 외

Recently, significant progress has been made in learned image and video compression. In particular the usage of Generative Adversarial Networks has lead to impressive results in the low bit rate regime. However, the mode…

Knowledge DistillationVideo Compression