paper-with-me

Papers

Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models

2026-03-01 · Adel Javanmard, Baharan Mirzasoleiman, Vahab Mirrokni arxiv

Large Language Models (LLMs) are pretrained on massive datasets and later instruction-tuned via supervised fine-tuning (SFT) or reinforcement learning (RL). Best practices emphasize large, diverse pretraining data, whereas post-training operates differently: SFT relies on smaller, high-quality datasets, while RL benefits more from scale, with larger amounts of feedback often outweighing label quality. Yet it remains unclear why pretraining and RL require large datasets, why SFT excels on smaller ones, and what defines high-quality SFT data. In this work, we theoretically analyze transformers trained on an in-context weight prediction task for linear regression. Our analysis reveals several key findings: $(i)$ balanced pretraining data can induce latent capabilities later activated during post-training, and $(ii)$ SFT learns best from a small set of examples challenging for the pretrained model, while excessively large SFT datasets may dilute informative pretraining signals. In contrast, RL is most effective on large-scale data that is not overly difficult for the pretrained model. We validate these theoretical insights with experiments on large nonlinear transformer architectures.

📄 PDF Abstract BibTeX arXiv:2603.01293

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Explaining Synergistic Effects in Social Recommendations

2026-01-26 · Yicong Li, Shan Jin, Qi Liu, Shuo Wang 외 arxiv

In social recommenders, the inherent nonlinearity and opacity of synergistic effects across multiple social networks hinders users from understanding how diverse information is leveraged for recommendations, consequently…

Cooperative effects in feature importance of individual patterns: application to air pollutants and Alzheimer disease

2025-07-30 · M. Ontivero-Ortega, A. Fania, A. Lacalamita, R. Bellotti 외 arxiv

Leveraging recent advances in the analysis of synergy and redundancy in systems of random variables, an adaptive version of the widely used metric Leave One Covariate Out (LOCO) has been recently proposed to quantify coo…

Feature Importance

The quest to quantify selective and synergistic effects of plasma for cancer treatment: Insights from mathematical modeling

2021-05-11 · Charlotta Bengtson, Annemie Bogaerts

Cold atmospheric plasma (CAP) and plasma-treated liquids (PTLs) have recently become a promising option for cancer treatment, but the underlying mechanisms of the anti-cancer effect are still to a large extent unknown. A…

Synergistic Image and Feature Adaptation: Towards Cross-Modality Domain Adaptation for Medical Image Segmentation

2019-01-24 · Cheng Chen, Qi Dou, Hao Chen, Jing Qin 외

This paper presents a novel unsupervised domain adaptation framework, called Synergistic Image and Feature Adaptation (SIFA), to effectively tackle the problem of domain shift. Domain adaptation has become an important a…

Domain AdaptationImage SegmentationMedical Image SegmentationSemantic Segmentation+1

Unsupervised Bidirectional Cross-Modality Adaptation via Deeply Synergistic Image and Feature Alignment for Medical Image Segmentation

2020-02-06 · Cheng Chen, Qi Dou, Hao Chen, Jing Qin 외

Unsupervised domain adaptation has increasingly gained interest in medical image computing, aiming to tackle the performance degradation of deep neural networks when being deployed to unseen data with heterogeneous chara…

Domain AdaptationImage SegmentationMedical Image SegmentationOrgan Segmentation+3