paper-with-me

홈 › Papers

Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs

2026-04-03 · Giyeong Oh, Junghyun Lee, Jaehyun Park, Youngjae Yu, Wonho Bae, Junhyug Noh arxiv

Modern LLMs inherit strong priors from web-scale pretraining, which can limit the headroom of post-training data-selection strategies. While Active Preference Learning (APL) seeks to optimize query efficiency in online Direct Preference Optimization (DPO), the inherent richness of on-policy candidate pools often renders simple Random sampling a surprisingly formidable baseline. We evaluate uncertainty-based APL against Random across harmlessness, helpfulness, and instruction-following settings, utilizing both reward models and LLM-as-a-judge proxies. We find that APL yields negligible improvements in proxy win-rates compared to Random. Crucially, we observe a dissociation where win-rate improves even as general capability -- measured by standard benchmarks -- degrades. APL fails to mitigate this capability collapse or reduce variance significantly better than random sampling. Our findings suggest that in the regime of strong pre-trained priors, the computational overhead of active selection is difficult to justify against the ``cheap diversity'' provided by simple random samples. Our code is available at https://github.com/BootsofLagrangian/random-vs-apl.

📄 PDF Abstract BibTeX arXiv:2604.02766

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure

2026-06-24 · Maxwell Twelftree, David Lemphers, An-chi He, Yue Yang arxiv

DiLoCo-style training reduces communication by letting learner islands train locally before occasional outer synchronization, making it attractive for fragmented industrial AI fleets where training shares hardware with l…

Forecasting skill of a crowd-prediction platform: A comparison of exchange rate forecasts

2023-12-14 · Niklas Valentin Lehmann

Open online crowd-prediction platforms are increasingly used to forecast trends and complex events. Despite the large body of research on crowd-prediction and forecasting tournaments, online crowd-prediction platforms ha…

Prediction

Humanoid Robot Running Through Random Stepping Stones and Jumping Over Obstacles: Step Adaptation Using Spring-Mass Trajectories

2025-12-15 · Sait Sovukluk, Johannes Englsberger, Christian Ott arxiv

This study proposes a step adaptation framework for running through spring-mass trajectories and deadbeat control gain libraries. It includes four main parts: (1) Automatic spring-mass trajectory library generation; (2) …

Nonlinear dynamics and fluctuations in biological systems (Habilitation thesis)

2018-03-20

The present habilitation thesis in theoretical biological physics addresses two central dynamical processes in cells and organisms: (i) active motility and motility control and (ii) self-organized pattern formation. The …

The Art of Beating the Odds with Predictor-Guided Random Design Space Exploration

2025-02-25 · Felix Arnold, Maxence Bouvier, Ryan Amaudruz, Renzo Andri 외

This work introduces an innovative method for improving combinational digital circuits through random exploration in MIG-based synthesis. High-quality circuits are crucial for performance, power, and cost, making this a …

Prediction