paper-with-me

홈 › Papers

Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?

2025-05-27 · ZiMing Wang, Zeyu Shi, Haoyi Zhou, Shiqi Gao, Qingyun Sun, JianXin Li

Fine-tuned Large Language Models (LLMs) often demonstrate poor calibration, with their confidence scores misaligned with actual performance. While calibration has been extensively studied in models trained from scratch, the impact of LLMs' prior knowledge on calibration during fine-tuning remains understudied. Our research reveals that LLMs' prior knowledge causes potential poor calibration due to the ubiquitous presence of known data in real-world fine-tuning, which appears harmful for calibration. Specifically, data aligned with LLMs' prior knowledge would induce overconfidence, while new knowledge improves calibration. Our findings expose a tension: LLMs' encyclopedic knowledge, while enabling task versatility, undermines calibration through unavoidable knowledge overlaps. To address this, we propose CogCalib, a cognition-aware framework that applies targeted learning strategies according to the model's prior knowledge. Experiments across 7 tasks using 3 LLM families prove that CogCalib significantly improves calibration while maintaining performance, achieving an average 57\% reduction in ECE compared to standard fine-tuning in Llama3-8B. These improvements generalize well to out-of-domain tasks, enhancing the objectivity and reliability of domain-specific LLMs, and making them more trustworthy for critical human-AI interaction applications.

📄 PDF Abstract BibTeX arXiv:2505.20903

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs

2025-10-10 · Xu Pan, Ely Hahami, Jingxuan Fan, Ziqian Xie 외 arxiv

Large language models (LLMs) are often used in environments where facts evolve, yet factual knowledge updates via fine-tuning on unstructured text often suffer from 1) reliance on compute-heavy paraphrasing augmentation …

Diversity in Large Language Models under Supervised Fine-Tuning

2026-04-30 · Roman Klypa, Oleksandr Cherednichenko arxiv

Supervised Fine-Tuning (SFT) is essential for aligning Large Language Models (LLMs) with user intent, yet it is believed to suppress generative diversity. Although this reduction is frequently referenced, formal empirica…

Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning

2026-01-23 · Qinglong Cao, Yuntian Chen, Chao Ma, Xiaokang Yang arxiv

Multimodal large language models (MLLMs) have shown remarkable capabilities in multimodal perception and understanding tasks. However, their effectiveness in specialized domains, such as remote sensing and medical imagin…

Domain Adaptation

Retention analysis of edited knowledge after fine-tuning

2025-07-14 · Fufang Wen, Shichang Zhang arxiv

Large language models (LLMs) store vast amounts of knowledge, which often requires updates to correct factual errors, incorporate newly acquired information, or adapt model behavior. Model editing methods have emerged as…

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs

2026-05-09 · Sangyeon Yoon, Wonje Jeung, Yoonjun Cho, Dongjae Jeon 외 arxiv

Fine-tuning APIs make frontier LLMs easy to customize, but they can also weaken safety alignment during fine-tuning. While prior work shows that benign supervised fine-tuning (SFT) can reduce refusal behavior, deployed f…