paper-with-me

홈 › Papers

Demystifying Data Organization for Enhanced LLM Training

2026-05-28 · Yalun Dai, Yangyu Huang, Tongshen Yang, Yonghan Wang, Xin Zhang, Wenshan Wu, Qihao Zhao, Hao Li, Yuanyuan Gao, Kim-Hui Yap, Scarlett Li arxiv

Large Language Models (LLMs) have revolutionized various fields, yet their training efficiency is heavily reliant on effective data curation. While data selection has been widely studied, the strategic data organization for enhanced training remains an underexplored area, particularly since current LLMs are often trained for only one or a few epochs. This paper systematically explores the influence of data organization on LLM training by reusing pre-computed sample-level scores originally generated for data efficiency, thereby incurring minimal additional computational overhead. We identify and formalize four key guidelines for optimizing data organization: Boundary Sharpening, Cyclic Scheduling, Curriculum Continuity, and Local Diversity. Guided by them, we introduce two novel data ordering methods termed STR and SAW. Extensive experiments across different model scales and data sizes, encompassing both pre-training and SFT stages, validate the effectiveness of our summarized guidelines. They also demonstrate the robustness of our proposed data ordering methods in enhancing the stability and performance of LLM training. Github Link: https://github.com/microsoft/data-efficacy/

📄 PDF Abstract BibTeX arXiv:2605.30334

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

It’s Commonsense, isn’t it? Demystifying Human Evaluations in Commonsense-Enhanced NLG Systems

2021-04-01 · EACL (HumEval) 2021 4 · Miruna-Adriana Clinciu, Dimitra Gkatzia, Saad Mahamood

Common sense is an integral part of human cognition which allows us to make sound decisions, communicate effectively with others and interpret situations and utterances. Endowing AI systems with commonsense knowledge cap…

Common Sense ReasoningText Generation

Demystifying Flux Architecture

2025-07-13 · Or Greenberg arxiv

FLUX.1 is a diffusion-based text-to-image generation model developed by Black Forest Labs, designed to achieve faithful text-image alignment while maintaining high image quality and diversity. FLUX is considered state-of…

Text-to-Image Generation

Demystifying Brain Tumour Segmentation Networks: Interpretability and Uncertainty Analysis

2019-09-03 · Parth Natekar, Avinash Kori, Ganapathy Krishnamurthi

The accurate automatic segmentation of gliomas and its intra-tumoral structures is important not only for treatment planning but also for follow-up evaluations. Several methods based on 2D and 3D Deep Neural Networks (DN…

Brain Tumor SegmentationMedical DiagnosisSegmentationTumor Segmentation

Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe

2026-03-23 · Xixi Wu, Qianguo Sun, Ruiyang Zhang, Chao Song 외 arxiv

Reinforcement Learning (RL) is essential for evolving Large Language Models (LLMs) into autonomous agents capable of long-horizon planning, yet a practical recipe for scaling RL in complex, multi-turn environments remain…

Reinforcement Learning

Generative AI and Organizational Structure in the Knowledge Economy

2025-05-31 · Fasheng Xu, Jing Hou, Wei Chen, Karen Xie

The adoption of GenAI is fundamentally reshaping organizations in the knowledge economy. GenAI can significantly enhance workers' problem-solving abilities and productivity, yet it also presents a major reliability chall…

Hallucination