paper-with-me

홈 › Papers

Using Small Proxy Datasets to Accelerate Hyperparameter Search

2019-06-12 · Sam Shleifer, Eric Prokop

One of the biggest bottlenecks in a machine learning workflow is waiting for models to train. Depending on the available computing resources, it can take days to weeks to train a neural network on a large dataset with many classes such as ImageNet. For researchers experimenting with new algorithmic approaches, this is impractically time consuming and costly. We aim to generate smaller "proxy datasets" where experiments are cheaper to run but results are highly correlated with experimental results on the full dataset. We generate these proxy datasets using by randomly sampling from examples or classes, training on only the easiest or hardest examples and training on synthetic examples generated by "data distillation". We compare these techniques to the more widely used baseline of training on the full dataset for fewer epochs. For each proxying strategy, we estimate three measures of "proxy quality": how much of the variance in experimental results on the full dataset can be explained by experimental results on the proxy dataset. Experiments on Imagenette and Imagewoof (Howard, 2019) show that running hyperparameter search on the easiest 10% of examples explains 81% of the variance in experiment results on the target task, and using the easiest 50% of examples can explain 95% of the variance, significantly more than training on all the data for fewer epochs, a more widely used baseline. These "easy" proxies are higher quality than training on the full dataset for a reduced number of epochs (but equivalent computational cost), and, unexpectedly, higher quality than proxies constructed from the hardest examples. Without access to a trained model, researchers can improve proxy quality by restricting the subset to fewer classes; proxies built on half the classes are higher quality than those with an equivalent number of examples spread across all classes.

📄 PDF Abstract BibTeX arXiv:1906.04887

Code (1)

DeepMindv2/DeepDetection tf

Similar Papers 제목 키워드 기반

Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice

2025-12-30 · Jiachen T. Wang, Tong Wu, Kaifeng Lyu, James Zou 외 arxiv

Data teams at frontier AI companies routinely train small proxy models to make critical decisions about pretraining data recipes for full-scale training runs. However, the community has a limited understanding of whether…

Hyperparameter Optimization

Hyperparameter Optimization with Neural Network Pruning

2022-05-18 · Kangil Lee, Junho Yim

Since the deep learning model is highly dependent on hyperparameters, hyperparameter optimization is essential in developing deep learning model-based applications, even if it takes a long time. As service development us…

Bayesian OptimizationDeep LearningHyperparameter OptimizationNetwork Pruning

Hyperparameter Transfer in Graph Neural Networks

2026-07-06 · Gage DeZoort, Boris Hanin arxiv

The performance of deep learning models crucially depends on the settings of hyperparameters like learning rate, initialization scale, and weight decay. Hyperparameter transfer aims to make near-optimal hyperparameter se…

Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler

2024-08-23 · Yikang Shen, Matthew Stallone, Mayank Mishra, Gaoyuan Zhang 외

Finding the optimal learning rate for language model pretraining is a challenging task. This is not only because there is a complicated correlation between learning rate, batch size, number of training tokens, model size…

ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization

2025-08-15 · Shengzhuang Chen, Xu Ouyang, Michael Arthur Leopold Pearce, Thomas Hartvigsen 외 arxiv

Determining the optimal data mixture for large language model training remains a challenging problem with an outsized impact on performance. In practice, language model developers continue to rely on heuristic exploratio…

Hyperparameter Optimization