paper-with-me

홈 › Papers

Domain Specific Specialization in Low-Resource Settings: The Efficacy of Offline Response-Based Knowledge Distillation in Large Language Models

2026-01-05 · Erdem Aslan, Pakize Erdoğmuş arxiv

Large Language Models (LLMs) excel in general tasks but often struggle with hallucinations when handling domain-specific or institutional knowledge absent from their pre-training. We present an offline response-based knowledge distillation method that develops high-accuracy specialized assistants under constrained hardware resources. We evaluate three distinct data strategies: general domain adaptation (15,000 lines), unstructured knowledge injection (2,000 lines), and a context-aware synthetic dataset (500 lines) generated by a teacher model. To minimize computational costs, we utilize the Unsloth library to optimize the Qwen-2.5-7B student model, reducing NVIDIA A100 GPU memory requirements from 40 GB to 16 GB. Experimental results demonstrate that while larger unstructured datasets suffer from persistent hallucinations, the 500-line context-aware dataset achieves a 96.7% accuracy rate and robust rejection capability. These findings validate the LIMA hypothesis, showing that data quality and structural alignment are more critical than quantity for domain adaptation in low-resource settings.

📄 PDF Abstract BibTeX arXiv:2601.16219

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationDomain Adaptation

Similar Papers 제목 키워드 기반

DS-TOD: Efficient Domain Specialization for Task-Oriented Dialog

2022-05-01 · Findings (ACL) 2022 5 · Chia-Chien Hung, Anne Lauscher, Simone Ponzetto, Goran Glavaš

Recent work has shown that self-supervised dialog-specific pretraining on large conversational datasets yields substantial gains over traditional language modeling (LM) pretraining in downstream task-oriented dialog (TOD…

dialog state trackingLanguage ModelingLanguage ModellingMasked Language Modeling+1

DS-TOD: Efficient Domain Specialization for Task Oriented Dialog

2021-10-15 · Chia-Chien Hung, Anne Lauscher, Simone Paolo Ponzetto, Goran Glavaš

Recent work has shown that self-supervised dialog-specific pretraining on large conversational datasets yields substantial gains over traditional language modeling (LM) pretraining in downstream task-oriented dialog (TOD…

dialog state trackingLanguage ModelingLanguage ModellingMasked Language Modeling+1

DS-TOD: Efficient Domain Specialization for Task-Oriented Dialog

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Recent work has shown that self-supervised dialog-specific pretraining on large conversational datasets yields substantial gains over traditional language modeling (LM) pretraining in downstream task-oriented dialog (TOD…

dialog state trackingLanguage ModelingLanguage ModellingMasked Language Modeling+1

DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs

2026-05-18 · Jing Wang, Hongxuan Lu, Jazze Young, Shu Wang 외 arxiv

Expert specialization in Mixture-of-Experts (MoE) models remains poorly understood, with traditional evaluations conflating architectural load-balancing with functional specialization. We introduce DBES, a comprehensive …

MoLoRA: Composable Specialization via Per-Token Adapter Routing

2026-03-16 · Shrey Shah, Justin Wagle arxiv

Multi-adapter serving systems route entire sequences to a single adapter, forcing a choice when requests span multiple domains. This assumption fails in two important settings: (1) multimodal generation, where text and i…

multimodal generation