LMEMs for post-hoc analysis of HPO Benchmarking
The importance of tuning hyperparameters in Machine Learning (ML) and Deep Learning (DL) is established through empirical research and applications, evident from the increase in new hyperparameter optimization (HPO) algorithms and benchmarks steadily added by the community. However, current benchmarking practices using averaged performance across many datasets may obscure key differences between HPO methods, especially for pairwise comparisons. In this work, we apply Linear Mixed-Effect Models-based (LMEMs) significance testing for post-hoc analysis of HPO benchmarking runs. LMEMs allow flexible and expressive modeling on the entire experiment data, including information such as benchmark meta-features, offering deeper insights than current analysis practices. We demonstrate this through a case study on the PriorBand paper's experiment data to find insights not reported in the original work.
Code (1)
Tasks
BenchmarkingHyperparameter OptimizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Towards Inferential Reproducibility of Machine Learning Research
Reliability of machine learning evaluation -- the consistency of observed evaluation scores across replicated model training runs -- is affected by several sources of nondeterminism which can be regarded as measurement n…
FairX: A comprehensive benchmarking tool for model analysis using fairness, utility, and explainability
We present FairX, an open-source Python-based benchmarking tool designed for the comprehensive analysis of models under the umbrella of fairness, utility, and eXplainability (XAI). FairX enables users to train benchmarki…
BenchmarkingFairnessPost-FEC BER Benchmarking for Bit-Interleaved Coded Modulation with Probabilistic Shaping
Accurate performance benchmarking after forward error correction (FEC) decoding is essential for system design in optical fiber communications. Generalized mutual information (GMI) has been shown to be successful at benc…
BenchmarkingWhat Motivates You? Benchmarking Automatic Detection of Basic Needs from Short Posts
According to the self-determination theory, the levels of satisfaction of three basic needs (competence, autonomy and relatedness) have implications on people{'}s everyday life and career. We benchmark the novel task of …
BenchmarkingBinary ClassificationClassificationReal-World Blur Dataset for Learning and Benchmarking Deblurring Algorithms
Numerous learning-based approaches to single image deblurring for camera and object motion blurs have recently been proposed. To generalize such approaches to real-world blurs, large datasets of real blurred images and t…
BenchmarkingDeblurringImage DeblurringSingle Image Deblurring