paper-with-me

Papers

LMEMs for post-hoc analysis of HPO Benchmarking

2024-08-05 · Anton Geburek, Neeratyoy Mallik, Danny Stoll, Xavier Bouthillier, Frank Hutter

The importance of tuning hyperparameters in Machine Learning (ML) and Deep Learning (DL) is established through empirical research and applications, evident from the increase in new hyperparameter optimization (HPO) algorithms and benchmarks steadily added by the community. However, current benchmarking practices using averaged performance across many datasets may obscure key differences between HPO methods, especially for pairwise comparisons. In this work, we apply Linear Mixed-Effect Models-based (LMEMs) significance testing for post-hoc analysis of HPO benchmarking runs. LMEMs allow flexible and expressive modeling on the entire experiment data, including information such as benchmark meta-features, offering deeper insights than current analysis practices. We demonstrate this through a case study on the PriorBand paper's experiment data to find insights not reported in the original work.

📄 PDF Abstract BibTeX arXiv:2408.02533

Code (1)

automl/lmem-significance 공식 구현

Tasks

BenchmarkingHyperparameter Optimization

Methods 이 논문이 사용한 방법론

HPO In machine learning, a hyperparameter is a parameter whose value is used to control learning process, and HPO is the problem of choosing a set of optimal hyperparameters for a…

Similar Papers 제목 키워드 기반

Towards Inferential Reproducibility of Machine Learning Research

2023-02-08 · Michael Hagmann, Philipp Meier, Stefan Riezler

Reliability of machine learning evaluation -- the consistency of observed evaluation scores across replicated model training runs -- is affected by several sources of nondeterminism which can be regarded as measurement n…

FairX: A comprehensive benchmarking tool for model analysis using fairness, utility, and explainability

2024-06-20 · Md Fahim Sikder, Resmi Ramachandranpillai, Daniel de Leng, Fredrik Heintz

We present FairX, an open-source Python-based benchmarking tool designed for the comprehensive analysis of models under the umbrella of fairness, utility, and eXplainability (XAI). FairX enables users to train benchmarki…

BenchmarkingFairness

Post-FEC BER Benchmarking for Bit-Interleaved Coded Modulation with Probabilistic Shaping

2020-04-24

Accurate performance benchmarking after forward error correction (FEC) decoding is essential for system design in optical fiber communications. Generalized mutual information (GMI) has been shown to be successful at benc…

Benchmarking

What Motivates You? Benchmarking Automatic Detection of Basic Needs from Short Posts

2021-08-01 · ACL 2021 5 · Sanja Stajner, Seren Yenikent, Bilal Ghanem, Marc Franco-Salvador

According to the self-determination theory, the levels of satisfaction of three basic needs (competence, autonomy and relatedness) have implications on people{'}s everyday life and career. We benchmark the novel task of …

BenchmarkingBinary ClassificationClassification

Real-World Blur Dataset for Learning and Benchmarking Deblurring Algorithms

2020-08-01 · ECCV 2020 8 · Jaesung Rim, Haeyun Lee, Jucheol Won, Sunghyun Cho

Numerous learning-based approaches to single image deblurring for camera and object motion blurs have recently been proposed. To generalize such approaches to real-world blurs, large datasets of real blurred images and t…

BenchmarkingDeblurringImage DeblurringSingle Image Deblurring