paper-with-me

홈 › Papers

The Hidden Costs of Translation Accuracy: Distillation, Quantization, and Environmental Impact

2025-09-28 · Dhaathri Vijay, Anandaswarup Vadapalli arxiv

The rapid expansion of large language models (LLMs) has heightened concerns about their computational and environmental costs. This study investigates the trade-offs between translation quality and efficiency by comparing full-scale, distilled, and quantized models using machine translation as a case study. We evaluated performance on the Flores+ benchmark and through human judgments of conversational translations in French, Hindi, and Kannada. Our analysis revealed that the full 3.3B FP32 model, while achieving the highest BLEU scores, incurred the largest environmental footprint (~ 0.007-0.008 kg CO2 per run). The distilled 600M FP32 model reduced inference time by 71-78% and carbon emissions by 63-65% compared with the full model, with only minimal reductions in BLEU scores. Human evaluations further showed that even aggressive quantization (INT4) preserved high levels of accuracy and fluency, with differences between models generally minor. These findings demonstrate that model compression strategies can substantially reduce computational demands and environmental impact while maintaining competitive translation quality, though trade-offs are more pronounced in low-resource settings. We argue for evaluation frameworks that integrate efficiency and sustainability alongside accuracy as central dimensions of progress in NLP.

📄 PDF Abstract BibTeX arXiv:2509.23990

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationModel Compression

Similar Papers 제목 키워드 기반

DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition

2023-05-18 · Hang Shao, Bei Liu, Wei Wang, Xun Gong 외

As a popular multilingual and multitask pre-trained speech model, Whisper has the problem of curse of multilinguality. To enhance multilingual capabilities in small Whisper models, we propose DQ-Whisper, a novel joint di…

Knowledge DistillationQuantizationspeech-recognitionSpeech Recognition

Model compression via distillation and quantization

2018-02-15 · ICLR 2018 1 · Antonio Polino, Razvan Pascanu, Dan Alistarh

Deep neural networks (DNNs) continue to make significant advances, solving tasks from image classification to translation or reinforcement learning. One aspect of the field receiving considerable attention is efficiently…

image-classificationmodelModel CompressionQuantization+1

Towards Full Utilization on Mask Task for Distilling PLMs into NMT

2021-09-17 · ACL ARR September 2021 9 · Anonymous

Owing to being well-performed in many natural language processing tasks, the application of pre-trained language models (PLMs) in neural machine translation (NMT) is widely concerned. Knowledge distillation (KD) is one o…

Knowledge DistillationMachine TranslationNMTTranslation

Lost in Distillation: A Case Study in Toxicity Modeling

2022-07-01 · NAACL (WOAH) 2022 7 · Alyssa Chvasta, Alyssa Lees, Jeffrey Sorensen, Lucy Vasserman 외

In an era of increasingly large pre-trained language models, knowledge distillation is a powerful tool for transferring information from a large model to a smaller one. In particular, distillation is of tremendous benefi…

Knowledge Distillation

Pieces of Eight: 8-bit Neural Machine Translation

2018-04-13 · NAACL 2018 6 · Jerry Quinn, Miguel Ballesteros

Neural machine translation has achieved levels of fluency and adequacy that would have been surprising a short time ago. Output quality is extremely relevant for industry purposes, however it is equally important to prod…

Machine TranslationQuantizationTranslation