paper-with-me

홈 › Papers

FRAMES: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy

2025-02-08 · Xuemiao Zhang, Feiyu Duan, Liangyu Xu, Yongwei Zhou, Sirui Wang, Rongxiang Weng, Jingang Wang, Xunliang Cai

Large language models (LLMs) have significantly advanced human language understanding and generation, with pretraining data quality and organization being crucial to their performance. Multi-stage pretraining is a promising approach, but existing methods often lack quantitative criteria for data partitioning and instead rely on intuitive heuristics. In this paper, we propose the novel Four-quadRAnt Multi-stage prEtraining strategy (FRAME), guided by the established principle of organizing the pretraining process into four stages to achieve significant loss reductions four times. This principle is grounded in two key findings: first, training on high Perplexity (PPL) data followed by low PPL data, and second, training on low PPL difference (PD) data followed by high PD data, both causing the loss to drop significantly twice and performance enhancements. By partitioning data into four quadrants and strategically organizing them, FRAME achieves a remarkable 16.8% average improvement over random across MMLU and CMMLU for the 3B model, effectively boosting LLM performance.

📄 PDF Abstract BibTeX arXiv:2502.05551

Code (0)

등록된 구현이 없습니다.

Tasks

MMLU

Similar Papers 제목 키워드 기반

How Vocabulary Sharing Facilitates Multilingualism in LLaMA?

2023-11-15 · Fei Yuan, Shuai Yuan, Zhiyong Wu, Lei LI

Large Language Models (LLMs), often show strong performance on English tasks, while exhibiting limitations on other languages. What is an LLM's multilingual capability when it is trained only on certain languages? The un…

Qffusion: Controllable Portrait Video Editing via Quadrant-Grid Attention Learning

2025-01-11 · Maomao Li, Lijian Lin, Yunfei Liu, Ye Zhu 외

This paper presents Qffusion, a dual-frame-guided framework for portrait video editing. Specifically, we consider a design principle of ``animation for editing'', and train Qffusion as a general animation framework from …

Video EditingVideo Generation

All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training

2026-01-07 · Chi Liu, Xin Chen arxiv

Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs). However, GRPO inherits PPO's token-level clipping while replacing token-level adv…

Reinforcement LearningMathematical Reasoning

Systematizing LLM Persona Design: A Four-Quadrant Technical Taxonomy for AI Companion Applications

2025-11-04 · Esther Sun, Zichu Wu arxiv

The design and application of LLM-based personas in AI companionship is a rapidly expanding but fragmented field, spanning from virtual emotional companions and game NPCs to embodied functional robots. This diversity in …

ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment

2026-07-22 · Stefanos Gkikas, Yu Fang, Christian Arzate Cruz, Muhammad Umar Khan 외 arxiv

Automatic pain assessment from facial video remains challenging due to the spatial heterogeneity of pain-related facial cues. This study proposes ReFace, a spatial reorganization pipeline that divides facial input into f…