paper-with-me

홈 › Papers

BLADE: Scalable Bi-level Adaptive Data Selection for LLM Training

2026-06-17 · Jiaxing Wang, Deping Xiang, Jin Xu, Zirui Liu, Zicheng Zhang, Guoqiang Gong, Jun Fang, Chao Liu, Pengzhang Liu, Tongxuan Liu, Ke Zhang, Qixia Jiang arxiv

As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive learning trajectories. Beyond static heuristic filtering, advanced data selection methods for LLM training largely follow two paradigms, each with fundamental limitations. Influence-based methods provide principled bi-level objectives but require intractable inverse-Hessian computations, while excess-loss methods are computationally efficient but rely on a static reference model that becomes misaligned with the evolving proxy model during training. We propose BLADE (Bi-Level Adaptive Data sElection), a Hessian-free framework for data selection. BLADE reformulates the bi-level optimization problem underlying influence-based methods as a penalized single-level objective via Lagrange multipliers, avoiding inverse-Hessian computation while revealing a principled connection to excess-loss based data selection. The resulting objective recovers an excess-loss form but replaces the static reference model with a dynamic one that stays synchronized with training. Theoretically, we prove that this penalized formulation guarantees first-order convergence. For efficient online batch selection, we instantiate BLADE as a memoryless randomized block-coordinate Frank-Wolfe algorithm. Extensive experiments show that BLADE consistently outperforms state-of-the-art data selection baselines, providing a practical recipe for LLM training.

📄 PDF Abstract BibTeX arXiv:2606.18650

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference

2026-04-29 · Bodon Jeong, Hongsu Byun, Youngjae Kim, Weikuan Yu 외 arxiv

The increasing deployment of Large Language Model (LLM) inference on edge AI systems demands efficient execution under tight memory budgets. A key challenge arises from Key-Value (KV) caches, which often exceed available…

Unsupervised Modular Adaptive Region Growing and RegionMix Classification for Wind Turbine Segmentation

2026-01-07 · Raül Pérez-Gonzalo, Riccardo Magro, Andreas Espersen, Antonio Agudo arxiv

Reliable operation of wind turbines requires frequent inspections, as even minor surface damages can degrade aerodynamic performance, reduce energy output, and accelerate blade wear. Central to automating these inspectio…

BLADE: Filter Learning for General Purpose Computational Photography

2017-11-29 · Pascal Getreuer, Ignacio Garcia-Dorado, John Isidoro, Sungjoon Choi 외

The Rapid and Accurate Image Super Resolution (RAISR) method of Romano, Isidoro, and Milanfar is a computationally efficient image upscaling method using a trained set of filters. We describe a generalization of RAISR, w…

DemosaickingDenoisingImage Super-ResolutionSuper-Resolution

BladeYOLO: Wind Turbine Blade Defect Detection with Limited Annotations and Weak-Saliency Awareness

2026-07-30 · Yabin Xu, Fangtao Zhang, Fan Wang, Zhan Wang 외 arxiv

Wind turbine blade defect detection remains highly challenging in real-world inspection scenarios due to limited on-site data and the subtle visual characteristics of defects. In practice, blade defects are often small-s…

Non-iterative generation of an optimal mesh for a blade passage using deep reinforcement learning

2022-09-08 · Innyoung Kim, Sejin Kim, Donghyun You

A method using deep reinforcement learning (DRL) to non-iteratively generate an optimal mesh for an arbitrary blade passage is developed. Despite automation in mesh generation using either an empirical approach or an opt…

Computational EfficiencyDeep Reinforcement Learningreinforcement-learningReinforcement Learning (RL)