paper-with-me

홈 › Papers

NAS-Bench-360: Benchmarking Neural Architecture Search on Diverse Tasks

2021-10-12 · Renbo Tu, Nicholas Roberts, Mikhail Khodak, Junhong Shen, Frederic Sala, Ameet Talwalkar

Most existing neural architecture search (NAS) benchmarks and algorithms prioritize well-studied tasks, e.g. image classification on CIFAR or ImageNet. This makes the performance of NAS approaches in more diverse areas poorly understood. In this paper, we present NAS-Bench-360, a benchmark suite to evaluate methods on domains beyond those traditionally studied in architecture search, and use it to address the following question: do state-of-the-art NAS methods perform well on diverse tasks? To construct the benchmark, we curate ten tasks spanning a diverse array of application domains, dataset sizes, problem dimensionalities, and learning objectives. Each task is carefully chosen to interoperate with modern CNN-based search methods while possibly being far-afield from its original development domain. To speed up and reduce the cost of NAS research, for two of the tasks we release the precomputed performance of 15,625 architectures comprising a standard CNN search space. Experimentally, we show the need for more robust NAS evaluation of the kind NAS-Bench-360 enables by showing that several modern NAS procedures perform inconsistently across the ten tasks, with many catastrophically poor results. We also demonstrate how NAS-Bench-360 and its associated precomputed results will enable future scientific discoveries by testing whether several recent hypotheses promoted in the NAS literature hold on diverse tasks. NAS-Bench-360 is hosted at https://nb360.ml.cmu.edu.

📄 PDF Abstract BibTeX arXiv:2110.05668

Code (1)

rtu715/nas-bench-360 공식 구현 pytorch

Tasks

Benchmarkingimage-classificationImage ClassificationNeural Architecture Search

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

SemSegBench & DetecBench: Benchmarking Reliability and Generalization Beyond Classification

2025-05-23 · Shashank Agnihotri, David Schader, Jonas Jakubassa, Nico Sharei 외

Reliability and generalization in deep learning are predominantly studied in the context of image classification. Yet, real-world applications in safety-critical domains involve a broader set of semantic tasks, such as s…

BenchmarkingClassificationimage-classificationImage Classification+4

TopoBench: A Framework for Benchmarking Topological Deep Learning

2024-06-09 · Lev Telyatnikov, Guillermo Bernardez, Marco Montagna, Mustafa Hajij 외

This work introduces TopoBench, an open-source library designed to standardize benchmarking and accelerate research in topological deep learning (TDL). TopoBench decomposes TDL into a sequence of independent modules for …

BenchmarkingDeep Learning

NAS-Bench-101: Towards Reproducible Neural Architecture Search

2019-02-25 · Chris Ying, Aaron Klein, Esteban Real, Eric Christiansen 외

Recent advances in neural architecture search (NAS) demand tremendous computational resources, which makes it difficult to reproduce experiments and imposes a barrier-to-entry to researchers without access to large-scale…

BenchmarkingNeural Architecture Search

KernelFoundry: Hardware-aware evolutionary GPU kernel optimization

2026-03-12 · Nina Wiedemann, Quentin Leboutet, Michael Paulitsch, Diana Wofk 외 arxiv

Optimizing GPU kernels presents a significantly greater challenge for large language models (LLMs) than standard code generation tasks, as it requires understanding hardware architecture, parallel optimization strategies…

Code Generation

STRABLE: Benchmarking Tabular Machine Learning with Strings

2026-05-12 · Gioia Blayer, Myung Jun Kim, Félix Lefebvre, Lennart Purucker 외 arxiv

Benchmarking tabular learning has revealed the benefit of dedicated architectures, pushing the state of the art. But real-world tables often contain string entries, beyond numbers, and these settings have been understudi…