paper-with-me

홈 › Papers

MaP: A Unified Framework for Reliable Evaluation of Pre-training Dynamics

2025-10-10 · Jiapeng Wang, Changxin Tian, Kunlong Chen, Ziqi Liu, Jiaxin Mao, Wayne Xin Zhao, Zhiqiang Zhang, Jun Zhou arxiv

Reliable evaluation is fundamental to the progress of Large Language Models (LLMs), yet the evaluation process during pre-training is plagued by significant instability that obscures true learning dynamics. In this work, we systematically diagnose this instability, attributing it to two distinct sources: \textit{Parameter Instability} from training stochasticity and \textit{Evaluation Instability} from noisy measurement protocols. To counteract both sources of noise, we introduce \textbf{MaP}, a dual-pronged framework that synergistically integrates checkpoint \underline{M}erging \underline{a}nd the \underline{P}ass@k metric. Checkpoint merging smooths the parameter space by averaging recent model weights, while Pass@k provides a robust, low-variance statistical estimate of model capability. Extensive experiments show that MaP yields significantly smoother performance curves, reduces inter-run variance, and ensures more consistent model rankings. Ultimately, MaP provides a more reliable and faithful lens for observing LLM training dynamics, laying a crucial empirical foundation for LLM research.

📄 PDF Abstract BibTeX arXiv:2510.09295

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models

2025-11-20 · Yang Luo, Xuanlei Zhao, Baijiong Lin, Lingting Zhu 외 arxiv

Recent progress in generative video models, such as Veo-3, has shown surprising zero-shot reasoning abilities, creating a growing need for systematic and reliable evaluation. We introduce V-ReasonBench, a benchmark desig…

Video Generation

Word Sense Disambiguation: A Unified Evaluation Framework and Empirical Comparison

2017-04-01 · EACL 2017 4 · Aless Raganato, ro, Jose Camacho-Collados, Roberto Navigli

Word Sense Disambiguation is a long-standing task in Natural Language Processing, lying at the core of human language understanding. However, the evaluation of automatic systems has been problematic, mainly due to the la…

Word Sense Disambiguation

UniKnow: A Unified Framework for Reliable Language Model Behavior across Parametric and External Knowledge

2025-02-19 · Youna Kim, Hyuhng Joon Kim, Minjoon Choi, Sungmin Cho 외

Language models often benefit from external knowledge beyond parametric knowledge. While this combination enhances performance, achieving reliable knowledge utilization remains challenging, as it requires assessing the s…

InformativenessLanguage ModelingLanguage Modelling

Thresh: A Unified, Customizable and Deployable Platform for Fine-Grained Text Evaluation

2023-08-14 · David Heineman, Yao Dou, Wei Xu

Fine-grained, span-level human evaluation has emerged as a reliable and robust method for evaluating text generation tasks such as summarization, simplification, machine translation and news generation, and the derived a…

Machine TranslationMulti-Task LearningNews GenerationText Generation

SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

2026-04-04 · Guiyu Zhang, Yabo Chen, Xunzhi Xiang, Junchao Huang 외 arxiv

Controlling both camera motion and object dynamics is essential for coherent and expressive video generation, yet current methods typically handle only one motion type or rely on ambiguous 2D cues that entangle camera-in…

Video Generation